Network content classification method and device, equipment, storage medium and program product

By setting up a topic information database and matching detection mechanism in the network content classification, the problem of reduced accuracy caused by fluctuations in the semantic distribution of models is solved, timely identification of new topics and model updates are achieved, and the accuracy and efficiency of classification are improved.

CN120296167APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410009809.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the accuracy of the network content classification model decreases when the semantic distribution fluctuates, resulting in a decrease in classification accuracy.

Method used

Set up a topic information database, detect the target network content and topic information through matching. If the matching fails, obtain data samples and update the topic information database and classification model, generate high-frequency common substrings as new topic information, and update the classification model to identify new topics.

Benefits of technology

提高了网络内容分类的准确性,及时识别新话题并更新模型,减少了模型迭代误判风险,提升了分类效率和准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296167A_ABST
    Figure CN120296167A_ABST
Patent Text Reader

Abstract

The invention discloses a network content classification method and device, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: performing matching detection on target network content and topic information contained in a topic information base; in response to failure of matching between the target network content and any topic information in the topic information base, acquiring a data sample based on the target network content; the data sample is used for updating the topic information base and the first classification model; and in response to matching of the target network content with any topic information in the topic information base, processing the target network content through a first classification model to obtain a classification category of the target network content. According to the scheme, the classification accuracy of the network content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a network content classification method, device, equipment, storage medium, and program product. Background Art

[0002] Network content classification is an important task in Internet applications, which can be used to discover the category to which network content belongs, so as to provide support for subsequent processing of network content.

[0003] In the related art, network content classification can be achieved through a pre-set machine learning model. Specifically, developers pre-train a classification model through network content samples and the annotation information corresponding to the network content samples. During the application of the model, the network content is input into the classification model, and the classification category to which the network content belongs is output by the classification model.

[0004] However, since the semantic distribution of network content may fluctuate greatly within a certain period of time, the accuracy of the classification model decreases during this period, thus affecting the accuracy of network content classification. Summary of the Invention

[0005] Embodiments of this application provide a network content classification method, device, equipment, storage medium, and program product, which can improve the accuracy of abnormal detection of images. The technical solutions are as follows:

[0006] On the one hand, a network content classification method is provided, and the method includes:

[0007] Obtain target network content;

[0008] Perform a matching detection on the target network content and the topic information included in the topic information library; the topic information library includes the topic information corresponding to one or more topics respectively;

[0009] In response to the target network content matching any topic information in the topic information library, process the target network content through a first classification model to obtain the classification category of the target network content output by the first classification model;

[0010] In response to the target network content failing to match any topic information in the topic information library, obtain a data sample based on the target network content, and the data sample is used to update the topic information library and the first classification model.

[0011] On the other hand, a network content classification device is provided, and the device includes:

[0012] A content acquisition module, configured to obtain target network content;

[0013] A matching module, configured to perform matching detection on the target network content and the topic information included in the topic information library; the topic information library includes the topic information corresponding to one or more topics respectively.

[0014] A sample acquisition module, configured to, in response to the failure of the target network content to match any topic information in the topic information library, acquire a data sample based on the target network content; the data sample is used to update the topic information library and the first classification model.

[0015] A first classification module, configured to, in response to the matching of the target network content with any topic information in the topic information library, process the target network content through the first classification model to obtain the classification category of the target network content output by the first classification model.

[0016] In some embodiments, the sample acquisition module is configured to acquire a corresponding text sample based on the target network content, and the text sample is used to update the topic information library.

[0017] In some embodiments, the apparatus further includes:

[0018] A text sample acquisition module, configured to acquire multiple text samples.

[0019] A high-frequency substring extraction module, configured to extract one or more high-frequency common substrings from multiple text samples.

[0020] A topic information generation module, configured to generate new topic information based on one or more high-frequency common substrings.

[0021] A topic information addition module, configured to add the new topic information to the topic information library.

[0022] In some embodiments, the high-frequency substring extraction module is configured to

[0023] Obtain the occurrence frequency of each single character in multiple text samples.

[0024] Based on the occurrence frequency of each single character in multiple text samples, obtain high-frequency single characters in multiple text samples.

[0025] Extract each common substring with a specified length in multiple text samples.

[0026] Based on the high-frequency single characters in multiple text samples, traverse each text sample to obtain the word frequency of each common substring.

[0027] Obtain the high-frequency common substrings in multiple said text samples based on the word frequencies of each said common substring.

[0028] In some embodiments, the high-frequency substring extraction module is configured to increment the word frequency of the first common substring in the currently traversed text sample by 1 in response to any of the high-frequency single words being included in the first common substring.

[0029] In some embodiments, in response to the specified length including multiple lengths, the topic information generation module is configured to,

[0030] In response to the number of the high-frequency common substrings being multiple, merge the multiple high-frequency common substrings to obtain one or more merged common substrings;

[0031] Obtain the merged common substring as a new said topic information.

[0032] In some embodiments, the device further includes:

[0033] A sample classification acquisition module, configured to acquire the classification categories of the network content corresponding to each of the multiple said text samples; the classification category of the network content corresponding to the text sample is obtained by classifying and labeling the network content corresponding to the text sample;

[0034] A topic purity setting module, configured to set the topic purity of the new said topic information to pure color in response to the proportion of the number of text samples corresponding to the first classification category among the multiple said text samples being greater than the proportion threshold;

[0035] The topic purity setting module is further configured to set the topic purity of the new said topic information to variegated in response to the proportion of the number of text samples corresponding to any classification category among the multiple said text samples not being greater than the proportion threshold;

[0036] A topic purity writing module, configured to write the topic purity of the new said topic information corresponding to the new said topic information into the topic information library;

[0037] The first classification module is configured to, in response to the target network content matching any topic information in the topic information library and the topic purity of the topic information matching the target network content being variegated, process the target network content through the first classification model to obtain the classification category of the target network content.

[0038] In some embodiments, the device further includes:

[0039] The first category acquisition module is configured to, in response to the topic color of the new topic information being solid color, acquire the first classification category as the classification category of the new topic information;

[0040] The classification category writing module is configured to write the classification category of the new topic information corresponding to the new topic information into the topic information library;

[0041] The second category acquisition module is configured to, in response to the target network content matching any topic information in the topic information library and the topic color of the topic information matching the target network content being solid color, acquire the classification category of the topic information matching the target network content as the classification category of the target network content.

[0042] In some embodiments, the text sample acquisition module is configured to acquire multiple text samples generated within a specified time range before the current moment.

[0043] In some embodiments, the sample acquisition module is configured to,

[0044] acquire the classification category of the target network content obtained by classifying and labeling the target network content;

[0045] acquire the target network content and the classification category of the target network content as training samples; the training samples are used to update the parameters of the first classification model.

[0046] In some embodiments, the apparatus further includes: a model update module, configured to,

[0047] process the target network content through the first classification model to obtain a first classification result of the target network content, where the first classification result is used to indicate the probabilities of the target network content belonging to various classification categories;

[0048] obtain a loss function value based on the difference between the first classification result and the classification category of the target network content;

[0049] update the parameters of the first classification model based on the loss function value.

[0050] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the network content classification method as described in the embodiments of the present application above.

[0051] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the network content classification method as described in the embodiments of the present application above.

[0052] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the network content classification method described in the above embodiments.

[0053] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0054] A topic information library is set up, which contains topic information of each topic. When the target network content is received, first, the target network content is matched and detected with the topic information in the topic information library. If the target network content matches any topic information in the topic information library, it is considered that the target network content is probably the network content corresponding to a known topic. At this time, the target network content can be classified by a classification model. If the target network content fails to match any topic information in the topic information library, it is considered that the target network content is probably the network content corresponding to an unknown topic. Classifying the target network content by the classification model may not be accurate. At this time, data samples can be obtained from the target network content to update the above-mentioned topic information library and classification model, so that the topic information library can identify the network content of new topics in time, and the classification model can be updated as soon as possible, so as to accurately identify the network content corresponding to new topics, thereby improving the accuracy of network content classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0056] Figure 1 is a system composition diagram of a network content classification system involved in the present application;

[0057] Figure 2 is a flowchart of a network content classification method shown according to an exemplary embodiment;

[0058] Figure 3 is a schematic flowchart of a network content classification method shown according to an exemplary embodiment;

[0059] Figure 4 is a schematic flowchart of a network content classification method shown according to an exemplary embodiment;

[0060] Figure 5 is a schematic flowchart of a network content classification method shown according to an exemplary embodiment;

[0061] Figure 6 is a system framework diagram related to the present application;

[0062] Figure 7 is a line schematic diagram of the recall accuracy rate related to the present application;

[0063] Figure 8 is a structural block diagram of a network content classification device provided by an exemplary embodiment of the present application;

[0064] Figure 9 is a schematic structural diagram of a server provided by an exemplary embodiment of the present application. Detailed implementation manners

[0065] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0066] Before describing the various embodiments shown in the present application, several concepts related to the present application will be introduced first.

[0067] 1) AI (Artificial Intelligence)

[0068] AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0069] 2) ML (Machine Learning, machine learning)

[0070] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching.

[0071] 3) Cloud technology

[0072] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.

[0073] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.

[0074] 4) Artificial intelligence cloud service

[0075] The so-called artificial intelligence cloud service, generally also known as AIaaS (AI as a Service, which means "AI as a service" in Chinese). This is a current mainstream service mode of artificial intelligence platforms. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service mode is similar to opening an AI-themed mall: all developers can access and use one or more artificial intelligence services provided by the platform through the way of Application Programming Interface (API). Some senior developers can also use the AI frameworks and AI infrastructure provided by the platform to deploy and operate their own exclusive cloud artificial intelligence services.

[0076] 5) Topic

[0077] Content that triggers discussions, such as "The grand opening of the XX Games".

[0078] 6) Model prediction drift

[0079] In the actual Internet business scenarios, the machine learning model used for network content processing (such as network content classification) remains unchanged, but due to the change in the feature / semantic distribution of the input content, the indicators such as the accuracy rate and recall level of the model prediction become abnormal. The common inducement is the emergence of new topics in the network. However, as time goes by, the model prediction indicators will automatically return to normal. To overcome the errors during this period, a class of methods adopted is called model prediction drift.

[0080] Please refer to Figure 1 , which shows the system composition diagram of a network content classification system involved in various embodiments of the present application. As Figure 1 shown, the system includes a terminal 140 and a server 160; optionally, the system may further include a database 180.

[0081] The terminal 140 can be a terminal device with certain processing capabilities and interface display functions. For example, the terminal 140 can be a mobile phone, a tablet computer, an e-book reader, smart glasses, a laptop computer, a desktop computer, and so on.

[0082] The terminal 140 can be a terminal used by users or a terminal used by developers.

[0083] When the terminal 140 is implemented as a terminal used by developers, the developers can develop a machine learning model for network content classification through the terminal 140 and deploy the machine learning model to the server 160.

[0084] When the terminal 140 is implemented as a terminal used by a user, an application for uploading or obtaining network content may be installed in the terminal 140.

[0085] Among them, the server 160 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0086] Optionally, the above-mentioned server 160 may be a server that provides background services for the applications installed in the terminal 140. This background server may be for version management of the application, performing background processing on the images obtained by the application and returning the processing results, performing background training on the machine learning models developed by developers, and so on.

[0087] Or, the above-mentioned server 160 may also be used to provide network content classification services for other servers.

[0088] The above-mentioned database 180 may be a Redis database, or may also be other types of databases. Among them, the database 180 is used to store various types of data.

[0089] Optionally, the terminal 140 and the server 160 are connected through a communication network. Optionally, this communication network is a wired network or a wireless network.

[0090] Optionally, the system may further include a management device ( Figure 1 not shown), and the management device is connected to the server 160 through a communication network. Optionally, the communication network is a wired network or a wireless network.

[0091] Figure 2 is a flowchart of a network content classification method shown according to an exemplary embodiment. This method may be executed by a computer device. For example, this computer device may be a server, and this server may be the server 160 in the above Figure 1 shown embodiment. As Figure 2 shown, this network content classification method may include the following steps.

[0092] Step 210: Obtain target network content.

[0093] Among them, the above-mentioned target network content may be network content newly obtained by the computer device from the network.

[0094] The target network content may include content information of various modalities. For example, the target network content may include text, pictures, videos, audio, and so on.

[0095] Step 220: Perform a matching detection on the target network content and the topic information included in the topic information library; the topic information library includes topic information corresponding to one or more topics respectively.

[0096] In an embodiment of the present application, a topic information library may be preset in the computer device, and the topic information library may store topic information corresponding to various known topics.

[0097] After the computer device obtains the target network content, it may first perform a matching detection on the target network content and each piece of topic information in the topic information library. For example, perform a similarity detection on the target network content and each piece of topic information in the topic information library. If the similarity between the target network content and a certain piece of topic information in the topic information library is greater than the similarity threshold, it may be considered that the target network content matches the topic information.

[0098] Step 230: In response to the failure of the target network content to match any topic information in the topic information library, obtain a data sample based on the target network content; the data sample is used to update the topic information library and the first classification model.

[0099] Among them, the failure of the target network content to match any topic information in the topic information library may mean that the target network content fails to match (or does not match) all the topic information in the topic information library.

[0100] In an embodiment of the present application, when it is detected that the target network content matches any topic information in the topic information library, it may be considered that the target network content is probably the network content corresponding to an unknown topic. At this time, the classification result of the target network content may not be accurate. At this time, the classification model for network content classification may be updated so that the classification model can accurately classify the network content corresponding to newly emerging topics in a timely manner. At the same time, the topic information library is also updated so that the topic information library can match newly emerging topics as soon as possible.

[0101] Step 240: In response to the target network content matching any topic information in the topic information library, process the target network content through the first classification model to obtain the classification category of the target network content output by the first classification model.

[0102] In an embodiment of the present application, when it is detected that the target network content matches any topic information in the topic information library, it can be considered that the target network content is probably the network content corresponding to a known topic. At this time, the computer device can classify the target network content through a classification model.

[0103] In summary, the solution shown in the embodiment of the present application sets up a topic information library that contains topic information for each topic. When receiving the target network content, first, the target network content is matched and detected with the topic information in the topic information library. If the target network content matches any topic information in the topic information library, it is considered that the target network content is probably the network content corresponding to a known topic. At this time, the target network content can be classified through a classification model. If the target network content fails to match any topic information in the topic information library, it is considered that the target network content is probably the network content corresponding to an unknown topic. Classifying the target network content through a classification model may not be accurate. At this time, data samples can be obtained from the target network content to update the above-mentioned topic information library and classification model, so that the topic information library can identify the network content of new topics in a timely manner, and the classification model can be updated as soon as possible to accurately identify the network content corresponding to new topics, thereby improving the accuracy of network content classification.

[0104] Based on Figure 2 the embodiment shown, please refer to Figure 3 , which is a flowchart showing a method for classifying network content according to an exemplary embodiment. As Figure 3 shown, the above step 230 can be implemented as step 230a; the method may further include steps 260a to 260d.

[0105] Step 230a: Obtain a corresponding text sample based on the target network content, and this text sample is used to update the topic information library.

[0106] In an embodiment of the present application, the computer device can extract a text sample from the target network content.

[0107] For example, when the target network content contains text content, the computer device can extract an abstract of the text content to obtain a text sample corresponding to the text content; or, the computer device can use the text content contained in the target network content as the text sample corresponding to the text content; or, the computer device can also filter the text content contained in the target network content to obtain a text sample corresponding to the text content, and so on.

[0108] For another example, when the target network content includes a picture / video frame, the computer device can perform semantic recognition or character recognition on the picture / video frame to obtain the text content included in the picture / video frame, and then obtain a text sample based on the text content.

[0109] For another example, when the target network content includes audio, the computer device can perform speech recognition on the audio to obtain the text content corresponding to the audio, and then obtain a text sample based on the text content.

[0110] In the embodiment of the present application, the computer device can obtain a text sample through network content that fails to match all the topic information in the topic information library to update the topic information library, so as to provide a solution for updating the topic information library through text information and ensure the feasibility of updating the topic information library.

[0111] Step 260a: Obtain multiple text samples.

[0112] In a possible implementation manner, the computer device can obtain multiple text samples generated within a specified duration range before the current moment.

[0113] Among them, in order to ensure the timeliness and accuracy of discovering new topics in the network, the computer device can regularly obtain text samples generated in the most recent period of time before the current moment. For example, the computer device can obtain new text samples generated within 15 minutes before the current moment every 15 minutes.

[0114] Step 260b: Extract one or more high-frequency common substrings from multiple text samples.

[0115] Among them, the above-mentioned high-frequency common substring refers to a common substring with a relatively high occurrence frequency among multiple text samples. For example, the above-mentioned high-frequency common substring can refer to a common substring with an occurrence frequency higher than a certain occurrence frequency threshold among multiple text samples; or, the above-mentioned high-frequency common substring can also refer to one or more common substrings arranged in the front in descending order of occurrence frequency among multiple text samples.

[0116] In a possible implementation manner, the process of extracting high-frequency common substrings from multiple text samples may include:

[0117] Obtain the occurrence frequency of each single character in multiple text samples;

[0118] Based on the occurrence frequency of each single character in multiple text samples, obtain high-frequency single characters in multiple text samples;

[0119] Extract each common substring with a specified length in multiple text samples;

[0120] Based on the high-frequency single characters in multiple text samples, traverse each text sample to obtain the word frequencies of each common substring;

[0121] Based on the word frequencies of each common substring, obtain the high-frequency common substrings in multiple text samples.

[0122] Among them, the specified length can include one or more preset lengths. For each length, the computer device can determine the common substring corresponding to the length from multiple text samples.

[0123] In the embodiments of the present application, the computer device can combine the occurrence frequencies of each single character in multiple text samples to determine the word frequencies of each common substring in multiple text samples, so as to accurately extract the high-frequency common substrings from multiple text samples and ensure the accuracy of the extraction of high-frequency common substrings.

[0124] In a possible implementation manner, based on the high-frequency single characters in multiple text samples, traverse each text sample to obtain the word frequencies of each common substring, including:

[0125] In response to any high-frequency single character being included in the first common substring in the currently traversed text sample, increment the word frequency of the first common substring by 1.

[0126] In the embodiments of the present application, after the computer device is convinced of the common substrings in multiple text samples, it can traverse each text sample. During the process of traversing a text sample, for the common substring in the current text sample, if the common substring contains a high-frequency single character, increment the word frequency of the common substring by 1 until all text samples are traversed to obtain the word frequencies of each common substring. The above solution can reduce the complexity of the calculation process for counting high-frequency substrings.

[0127] Step 260c: Generate new topic information based on one or more high-frequency common substrings.

[0128] In the embodiments of the present application, after determining one or more high-frequency common substrings, new topic information can be generated based on the high-frequency common substrings, so as to discover new topics.

[0129] In a possible implementation manner, in response to the specified length including multiple lengths, the process of generating new topic information based on one or more high-frequency common substrings may include:

[0130] In response to the number of high-frequency common substrings being multiple, merge the multiple high-frequency common substrings to obtain one or more merged common substrings;

[0131] Obtain the merged common substring as a new topic information.

[0132] Among them, when the specified length includes multiple lengths, there may be a situation where a high-frequency common substring with a short length is included in a high-frequency common substring with a long length. In this case, the above-mentioned high-frequency common substring with a short length and the high-frequency common substring with a long length can be merged. Specifically, if a high-frequency common substring with a short length is included in a high-frequency common substring with a long length, then the high-frequency common substring with a short length is discarded, and the high-frequency common substring with a long length is retained as the merged common substring, so as to avoid duplication of topic information.

[0133] Step 260d: Add the new topic information to the topic information library.

[0134] In the embodiment of the present application, after generating the new topic information, the new topic information can be added to the topic information library, so as to filter known topics through the topic information library and avoid generating samples for subsequent network content corresponding to the new topic information.

[0135] In the embodiment of the present application, the computer device generates new topic information by extracting high-frequency common substrings from multiple text samples, thereby providing a solution for automatically discovering new topic information through statistical methods, ensuring the feasibility and accuracy of discovering new topic information.

[0136] Based on Figure 3 the embodiments shown, please refer to Figure 4 which is a schematic flowchart of a network content classification method shown according to an exemplary embodiment. As Figure 4 shown, the above solution may further include steps 260e to 260h; the above step 240 may be implemented as step 240a.

[0137] Step 260e: Obtain the classification categories of the network content corresponding to each of the multiple text samples; the classification category of the network content corresponding to the text sample is obtained by classifying and labeling the network content corresponding to the text sample.

[0138] In the embodiment of the present application, for the above multiple text samples, the computer device may also obtain the classification categories obtained by pre-classifying and labeling the network content corresponding to each of the multiple text samples. Among them, the above classification and labeling may refer to manual classification and labeling or semi-manual classification and labeling.

[0139] Among them, the above semi-manual classification and labeling may refer to that the computer device first preliminarily classifies the network content corresponding to each of the multiple text samples through a second classification model, and then manually corrects the results of the preliminary classification of the network content corresponding to each of the multiple text samples to obtain the classification categories obtained by classifying and labeling the network content corresponding to each of the multiple text samples.

[0140] Step 260f: In response to the proportion of text samples corresponding to the first classification category among multiple text samples being greater than the proportion threshold, set the topic color of the new topic information to solid color.

[0141] Among them, the above proportion threshold can be greater than or equal to 0.5. For example, the above proportion threshold can be set to 0.95, etc.

[0142] Among them, if most of the text samples among multiple text samples correspond to the same classification category, it means that the network content corresponding to the new topic information probably corresponds to the same classification category. At this time, the topic color of the topic information can be set to solid color.

[0143] Step 260g: In response to the proportion of text samples corresponding to any classification category among multiple text samples not being greater than the proportion threshold, set the topic color of the new topic information to variegated color.

[0144] Among them, if there is no situation where most of the text samples among multiple text samples correspond to the same classification category, it means that the network content corresponding to the new topic information may correspond to multiple different classification categories. At this time, the topic color of the topic information can be set to variegated color.

[0145] Step 260h: Write the new topic information corresponding to the topic color of the new topic information into the topic information library.

[0146] In the embodiments of the present application, when the computer device writes new topic information into the topic information library, it can also write the topic color of the new topic information correspondingly.

[0147] Step 240a: In response to the target network content matching any topic information in the topic information library and the topic color of the topic information matching the target network content being variegated color, process the target network content through the first classification model to obtain the classification category of the target network content.

[0148] In the embodiments of the present application, when the target network content matches any topic information in the topic information library, the computer device can further obtain the topic color of the topic information matching the target network content. If the obtained topic information is variegated color, it means that the classification category of the target network content may be one of multiple classification categories and cannot be directly determined. At this time, the target network content can be classified through the first classification model to ensure the accuracy of classification.

[0149] In a possible implementation manner, the method further includes:

[0150] In response to the topic color of the new topic information being solid color, obtain the first classification category as the classification category of the new topic information; write the classification category of the new topic information corresponding to the new topic information into the topic information library.

[0151] In response to the target network content matching any topic information in the topic information library and the topic color of the topic information matching the target network content being solid color, obtain the classification category of the topic information matching the target network content as the classification category of the target network content.

[0152] In the embodiment of the present application, if the topic color of the new topic information is solid color, then the network content corresponding to the new topic information in the network probably also belongs to the above-mentioned first classification category. At this time, the computer device can also set the first classification category as the classification category of the new topic information, and when writing the new topic information into the topic information library, also write the classification category of the new topic information into the topic information library; correspondingly, when the target network content matches any topic information in the topic information library, and the computer device obtains that the topic color of the topic information matching the target network content is solid color, it indicates that the classification category of the target network content is probably a certain fixed classification category, that is, the classification category corresponding to the topic information matching the target network content. At this time, the computer device can directly set the classification category of the target network content as the classification category corresponding to the topic information matching the target network content, without the need to process through the first classification model, thereby improving the classification efficiency.

[0153] Based on Figure 2 、 Figure 3 or Figure 4 shown in the embodiments, please refer to Figure 5 , which is a flowchart of a network content classification method shown according to an exemplary embodiment. As Figure 5 shown, the above step 230 can also be implemented as step 230b and step 230c.

[0154] Step 230b: Obtain the classification category of the target network content obtained by classifying and annotating the target network content.

[0155] Among them, the process of obtaining the classification category of the target network content obtained by classifying and annotating the target network content can refer to the process of obtaining the classification category of the network content corresponding to each of the above multiple text samples, which will not be elaborated here.

[0156] Step 230c: Obtain the target network content and the classification category of the target network content as training samples; the training samples are used to update the parameters of the first classification model.

[0157] In a possible implementation, the computer device may perform the following steps to update the parameters of the first classification model:

[0158] Process the target network content through the first classification model to obtain the first classification result of the target network content, where the first classification result is used to indicate the probabilities of the target network content belonging to various classification categories;

[0159] Obtain the loss function value based on the difference between the first classification result and the classification category of the target network content;

[0160] Update the parameters of the first classification model based on the loss function value.

[0161] In the embodiments of the present application, the computer device may construct training samples through network content that fails to match all the topic information in the topic information library to train the first classification model. Since the network content that fails to match all the topic information in the topic information library may be network content corresponding to a new topic, through the above solution, the acquisition and training of training samples corresponding to new topics can be automatically realized, so that the update of the first classification model is more targeted, thus ensuring the timeliness and accuracy of the update of the first classification model.

[0162] The solution shown in the above embodiments of the present application can be applied to various scenarios based on network content classification, such as recommendation scenarios, etc. The above solution provides the following:

[0163] 1) Fuzzy discovery of topics. Taking the text "The 19th xx Games in XX City grandly opened" as an example, for precise discovery in the recommendation scenario, it is necessary to identify content with complete semantics, such as content with a complete subject, adverbial, and predicate like "xx Games grandly opened", while the fuzzy discovery of topics involved in the present application only needs to identify "Games, grandly opened".

[0164] 2) Prediction offset of the network content recognition model based on the sudden topic discovered through fuzzy discovery. Specifically, a series of solutions based on fuzzy topic discovery are provided when a new topic occurs and the model prediction accuracy decreases, reducing the risk of misjudgment in model iteration, realizing automatic offset of the prediction result, greatly reducing the impact of the topic on the model, and reducing the pressure of model operation.

[0165] To better illustrate the implementation process of the above solution of this application, take the text multi-classification task of the text data stream service as an example. The classification labels are four categories: "Sports, Food, Games, Others". As time goes by, the text multi-classification model continuously makes model predictions and forms a certain data distribution within a period of time. For example, "The men's basketball final of XX Cup is about to start" belongs to the "Sports category", and "There are various kinds of pasta in XX" belongs to the "Food category". However, when the game "XX Elite" became extremely popular and its promotional slogan "Big X, big X, eat Y tonight" was widely spread, it was easily misidentified as the "Food category", but its correct category should be the "Games category". Through the above embodiments, please refer to Figure 6 , which shows the system framework diagram involved in this application, as Figure 6 shown, the solution process involved in this application can be as follows:

[0166] Step S1: The hot topic discovery method adopts "batch data topic discovery".

[0167] This application does not adopt the method of "single text topic discovery and then cumulative popularity". The main reason is that the correct type of the topic content after manual correction may also be incorrect. For example, "eating X" is not considered "games" by everyone, and it needs to be comprehensively judged based on the results of a group of people.

[0168] The embodiment of this application uses a time sliding window [t1, tn] to accumulate a sample set set(P1, P2,..., Pm) with a data scale of m to discover new topics. When a new topic comes, it is not recorded in the topic information database, so all data will flow into the normal model (the model for classifying network content). After accumulating data samples corrected manually for a period of time, it is prepared for fuzzy topic discovery.

[0169] Step S2: Fuzzy topic discovery.

[0170] The input is m batch texts, and the output is the set set(<g1, c1>; <g2, c2>... <gk, ck>), where g represents topic information and c represents topic quality. The specific algorithm steps are as follows:

[0171] S2.1, Topic information extraction.

[0172] Assume m = 3, and the batch texts are P1 = "Big X, big X, eat X tonight", P2 = "I'll wait for you to eat X tonight", P3 = "Call your friends and eat X tonight". It is necessary to extract the topic information g = "eat X tonight". The specific algorithm process is as follows:

[0173] a) Calculate the set A of hot topic atomic information.

[0174] A unit smaller than the topic information g is the topic atomic information a. In the text scenario, a is a text of length 1, that is, a single character, and a is a substring of s. The comparison cases of the two can be: "Eat X tonight" and "X". It can be understood that to obtain "Eat X tonight", only the text content before and after "X" needs to be concerned, because the hot topic information s must contain the hot topic atomic information a. The extraction algorithm is as follows:

[0175] Concatenate all m texts and count the frequencies of all single characters, extract the top-k high-frequency single characters, and define them as all a, that is, A = {No.1 high-frequency single character,..., No.k high-frequency single character}, and a ∈ A. For example, k = 30.

[0176] b) Traverse the m texts in a loop and count the common substrings g of length l for all texts. For example, l ∈ [4, 7]. The larger l is, the higher the topic extraction accuracy. Let the total number of g be e, then there is:

[0177]

[0178] e = (r - l + 1) * m

[0179] c) Calculate the set G of high-frequency common substrings g. During the process of traversing each text, if g contains any element in the set A, the word frequency of g is incremented by 1, otherwise it is directly skipped. The definition is as follows:

[0180] freq(g i ) = the word frequency of g (i ∈ [1, e]) i

[0181] This step can significantly optimize the overall algorithm time complexity. The time complexity optimization of this step is as follows:

[0182] Optimized from O(r.m) to O(|A|.m), where r >> |A|, and |A| is the number of elements in the set A, that is, k mentioned above.

[0183] d) Merge common substrings. For example, if g1 and g2 are both included in g3, then only g3 needs to be retained. For example, "Eat tonight", "Eat X tonight", and "Tonight eat X" are all high-frequency common substrings, then only "Tonight eat X" needs to be retained. The formula is expressed as:

[0184] G = set(g i , g j ), where i, j ∈ [1, e]

[0185] e) The output is the set of all topic information.

[0186] S2.2, extract the topic quality c.​

[0187] Among them, c ∈ [solid color, variegated color]. That is, there are 2 types of topic colors, namely solid color and variegated color. In this application, it can be defined that a solid-color topic is that the maximum class proportion of all sample types > 95%, otherwise it is a variegated-color topic. For example, after manual determination, the above P1, P2, and P3 all belong to the "game category", so the result of fuzzy topic discovery is: the set {<g = "What to eat tonight", c = "solid color">}, and it is added to the topic information library.

[0188] Step S3: Topic information matching.

[0189] As time goes by, the topic information has been recognized at time tn, and the new data "You said what to eat tonight, I thought you were going to invite me to eat XX" at time tn+1 enters the sample queue and matches successfully with {<g = "What to eat tonight", c = "solid color">} in the topic information library.

[0190] Step S4: Topic customization rule judgment.

[0191] According to the topic color, if the topic color = "solid color", for example, "What to eat tonight" is probably the "game category" in the short term and does not need to go through the topic model for content classification anymore. In addition, if the topic color = variegated color, it proves that the category is complex and specific categories need to be output according to the specific text content. At this time, a dedicated topic model (corresponding to the above first classification model) is required for classification. Here, various algorithms can be used for the topic model. For example, natural language processing (NLP) models can be used for text content, computer vision (CV) models can be used for picture content, and multimodal models can be used for multimodal content, etc.

[0192] Step S5: Topic model iteration.

[0193] Due to the rich and diverse topic content, for topics in different time periods, the topic model needs to be updated to achieve the classification effect. At this time, the topic model can be iterated through the topic data and data labels in the topic sample library.

[0194] The solution shown in the above embodiments of this application is more interpretable compared with the traditional method of directly classifying network content through a classification model and iterating the classification model regularly, and has obvious advantages in terms of indicators such as coverage rate and accuracy.

[0195] 1) Coverage rate index dimension

[0196] As shown in Table 1 below, the traditional method has the optimal prediction offset and the minimum coverage loss at a model score of 0.8. However, the solution provided in this application has a higher coverage rate for multiple labels than the traditional method.

[0197] Table 1

[0198]

[0199]

[0200] 2) Recall accuracy metric dimension

[0201] Please refer to Figure 7 , which shows the schematic diagram of the recall accuracy line related to this application. As Figure 7 shown, line 71 is the magnitude after the prediction offset of this application, line 72 is the hit magnitude under normal circumstances, line 73 is the magnitude after the prediction offset of the traditional method, and line 74 is the predicted magnitude when the topic occurs. In the case of the same coverage rate, the smaller the recall magnitude, the higher the accuracy. Here, the recall magnitude is equivalently used for comparison. When a hot topic occurs, using the solution shown in the embodiments of this application, a new topic can be accurately identified with less latency (15 minutes). That is to say, when a sudden hot topic occurs, the model prediction index of this application cannot return to the normal situation for a short period of time, that is, there is a gap D1 between line 71 and line 72. However, compared with the gap of the traditional method, that is, the gap D2 between line 73 and line 72, D1 << D2, which means that the solution of this application has a better effect.

[0202] Figure 8 is the structural block diagram of a network content classification device provided by an exemplary embodiment of this application. As Figure 8 shown, the device includes the following parts:

[0203] Content acquisition module 801, used to acquire target network content;

[0204] Matching module 802, used to perform matching detection on the target network content and the topic information included in the topic information library; the topic information library includes the topic information corresponding to one or more topics respectively;

[0205] Sample acquisition module 803, used to respond to the failure of the target network content to match any topic information in the topic information library, and acquire data samples based on the target network content; the data samples are used to update the topic information library and the first classification model;

[0206] The first classification module 804 is configured to, in response to the target network content matching any topic information in the topic information library, process the target network content through the first classification model, and obtain the classification category of the target network content output by the first classification model.

[0207] In some embodiments, the sample acquisition module 803 is configured to obtain a corresponding text sample based on the target network content, and the text sample is used to update the topic information library.

[0208] In some embodiments, the apparatus further includes:

[0209] A text sample acquisition module, configured to acquire multiple text samples;

[0210] A high-frequency substring extraction module, configured to extract one or more high-frequency common substrings from multiple text samples;

[0211] A topic information generation module, configured to generate new topic information based on one or more high-frequency common substrings;

[0212] A topic information addition module, configured to add the new topic information to the topic information library.

[0213] In some embodiments, the high-frequency substring extraction module is configured to,

[0214] Obtain the occurrence frequency of each single character in multiple text samples;

[0215] Based on the occurrence frequency of each single character in multiple text samples, obtain high-frequency single characters in multiple text samples;

[0216] Extract each common substring with a specified length in multiple text samples;

[0217] Based on the high-frequency single characters in multiple text samples, traverse each text sample to obtain the word frequency of each common substring;

[0218] Based on the word frequency of each common substring, obtain the high-frequency common substrings in multiple text samples.

[0219] In some embodiments, the high-frequency substring extraction module is configured to, in response to any of the high-frequency single characters being included in the first common substring in the currently traversed text sample, increment the word frequency of the first common substring by 1.

[0220] In some embodiments, in response to the specified length including multiple lengths, the topic information generation module is configured to,

[0221] In response to the number of the high-frequency common substrings being multiple, merge the multiple high-frequency common substrings to obtain one or more merged common substrings;

[0222] Take the merged common substring as a new piece of the topic information.

[0223] In some embodiments, the apparatus further includes:

[0224] A sample classification acquisition module, configured to acquire the classification categories of the network contents corresponding to multiple pieces of the text samples respectively; the classification categories of the network contents corresponding to the text samples are obtained by classifying and labeling the network contents corresponding to the text samples;

[0225] A topic purity setting module, configured to, in response to the proportion of the number of the text samples corresponding to the first classification category among the multiple pieces of the text samples being greater than a proportion threshold, set the topic purity of the new topic information to pure color;

[0226] The topic purity setting module is further configured to, in response to the proportion of the number of the text samples corresponding to any classification category among the multiple pieces of the text samples not being greater than the proportion threshold, set the topic purity of the new topic information to mixed color;

[0227] A topic purity writing module, configured to write the topic purity of the new topic information corresponding to the new topic information into the topic information library;

[0228] The first classification module 804 is configured to, in response to the target network content matching any topic information in the topic information library and the topic purity of the topic information matching the target network content being mixed color, process the target network content through the first classification model to obtain the classification category of the target network content.

[0229] In some embodiments, the apparatus further includes:

[0230] A first category acquisition module, configured to, in response to the topic purity of the new topic information being pure color, acquire the first classification category as the classification category of the new topic information;

[0231] A classification category writing module, configured to write the classification category of the new topic information corresponding to the new topic information into the topic information library;

[0232] The second category acquisition module is configured to, in response to the target network content matching any topic information in the topic information library and the topic color of the topic information matching the target network content being a solid color, acquire the classification category of the topic information matching the target network content as the classification category of the target network content.

[0233] In some embodiments, the text sample acquisition module is configured to acquire multiple text samples generated within a specified duration range before the current moment.

[0234] In some embodiments, the sample acquisition module 803 is configured to,

[0235] acquire the classification category of the target network content obtained by classifying and labeling the target network content;

[0236] acquire the target network content and the classification category of the target network content as training samples; the training samples are used to update the parameters of the first classification model.

[0237] In some embodiments, the apparatus further includes: a model update module, configured to,

[0238] process the target network content through the first classification model to obtain a first classification result of the target network content, where the first classification result is used to indicate the probabilities of the target network content belonging to various classification categories;

[0239] obtain a loss function value based on the difference between the first classification result and the classification category of the target network content;

[0240] update the parameters of the first classification model based on the loss function value.

[0241] In summary, the solution shown in the embodiments of the present application sets up a topic information library containing topic information for each topic. When receiving target network content, first, the target network content is matched and detected with the topic information in the topic information library. If the target network content matches any topic information in the topic information library, it is considered that the target network content is probably the network content corresponding to a known topic. At this time, the target network content can be classified through a classification model. If the target network content fails to match any topic information in the topic information library, it is considered that the target network content is probably the network content corresponding to an unknown topic. Classifying the target network content through the classification model may not be accurate. At this time, data samples can be obtained from the target network content to update the above-mentioned topic information library and classification model, so that the topic information library can identify the network content of new topics in a timely manner, and the classification model can be updated as soon as possible to accurately identify the network content corresponding to new topics, thereby improving the accuracy of network content classification.

[0242] It should be noted that: for the network content classification device provided in the above embodiments, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the network content classification device provided in the above embodiments and the embodiments of the network content classification method belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.

[0243] Figure 9 The structural schematic diagram of a server provided by an exemplary embodiment of the present application is shown. Specifically:

[0244] The server 900 includes a central processing unit (CPU) 901, a system memory 904 including a random access memory (RAM) 902 and a read-only memory (ROM) 903, and a system bus 905 connecting the system memory 904 and the central processing unit 901. The server 900 also includes a mass storage device 906 for storing an operating system 913, application programs 914, and other program modules 915.

[0245] The mass storage device 906 is connected to the central processing unit 901 through a mass storage controller (not shown) connected to the system bus 905. The mass storage device 906 and its associated computer-readable medium provide non-volatile storage for the server 900. That is, the mass storage device 906 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0246] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that computer storage media is not limited to the above several types. The above system memory 904 and mass storage device 906 can be collectively referred to as memory.

[0247] According to various embodiments of the present application, the server 900 can also run by connecting to a remote computer on the network through a network such as the Internet. That is, the server 900 can be connected to the network 912 through the network interface unit 911 connected to the system bus 905, or in other words, the network interface unit 911 can also be used to connect to other types of networks or remote computer systems (not shown).

[0248] The above memory further includes one or more programs, and one or more programs are stored in the memory and configured to be executed by the CPU.

[0249] Embodiments of the present application also provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, at least one segment of program, a code set or an instruction set, and the at least one instruction, at least one segment of program, code set or instruction set is loaded and executed by the processor to implement the network content classification method provided by the above method embodiments.

[0250] Embodiments of the present application also provide a computer-readable storage medium, on which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by a processor to implement the network content classification method provided in the above method embodiments.

[0251] Embodiments of the present application also provide a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the network content classification method described in any one of the above embodiments.

[0252] Optionally, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), solid state drive (SSD, Solid State Drives), or optical disc, etc. Among them, the random access memory may include resistive random access memory (ReRAM, Resistance RandomAccess Memory) and dynamic random access memory (DRAM, Dynamic Random Access Memory). The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0253] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be read-only memory, a magnetic disk or an optical disc, etc.

[0254] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for classifying network content, characterized in that, The method includes: Obtaining target network content; Performing matching detection on the target network content and the topic information included in a topic information library; the topic information library includes the topic information corresponding to one or more topics respectively; In response to the failure of the target network content to match any topic information in the topic information library, obtaining a data sample based on the target network content; the data sample is used to update the topic information library and a first classification model; In response to the target network content matching any topic information in the topic information library, processing the target network content through the first classification model to obtain the classification category of the target network content output by the first classification model.

2. The method according to claim 1, wherein The step of, in response to the failure of the target network content to match any topic information in the topic information library, obtaining a data sample based on the target network content includes: Obtaining a corresponding text sample based on the target network content, and the text sample is used to update the topic information library.

3. The method according to claim 2, characterized in that, The method further includes: Obtaining multiple such text samples; Extracting one or more high-frequency common substrings from the multiple text samples; Generating new topic information based on the one or more high-frequency common substrings; Adding the new topic information to the topic information library.

4. The method according to claim 3, wherein The step of extracting high-frequency common substrings from the multiple text samples includes: Obtaining the occurrence frequency of each single character in the multiple text samples; Based on the occurrence frequency of each single character in the multiple text samples, obtaining high-frequency single characters in the multiple text samples; Extracting each common substring with a specified length in the multiple text samples; Based on the high-frequency single characters in the multiple text samples, traversing each text sample to obtain the word frequency of each common substring; Based on the word frequency of each common substring, obtaining the high-frequency common substrings in the multiple text samples.

5. The method according to claim 4, characterized in that The step of, based on the high-frequency single characters in the multiple text samples, traversing each text sample to obtain the word frequency of each common substring includes: In response to the first common substring in the currently traversed text sample containing any of the high-frequency single characters, incrementing the word frequency of the first common substring by 1.

6. The method according to claim 3, wherein In response to the specified length including multiple lengths, the step of generating new topic information based on the one or more high-frequency common substrings includes: In response to the number of the high-frequency common substrings being multiple, merging the multiple high-frequency common substrings to obtain one or more merged common substrings; Taking the merged common substring as a new piece of topic information.

7. The method according to claim 3, characterized in that The method further includes: Obtaining the classification category of the network content corresponding to each of the multiple text samples; the classification category of the network content corresponding to the text sample is obtained by classifying and labeling the network content corresponding to the text sample; In response to the proportion of the number of text samples corresponding to the first classification category among the multiple text samples being greater than a proportion threshold, setting the topic quality of the new topic information to pure color. In response to the proportion of the number of text samples corresponding to any one classification category among the multiple text samples not being greater than the proportion threshold, set the topic color of the new topic information to variegated; Write the new topic information corresponding to the topic color of the new topic information into the topic information library; The response that the target network content matches any topic information in the topic information library, and processes the target network content through the first classification model to obtain the classification category of the target network content output by the first classification model, includes: In response to the target network content matching any topic information in the topic information library and the topic color of the topic information matching the target network content being variegated, process the target network content through the first classification model to obtain the classification category of the target network content.

8. The method according to claim 7, wherein The method further includes: In response to the topic color of the new topic information being pure color, obtain the first classification category as the classification category of the new topic information; Write the classification category of the new topic information corresponding to the new topic information into the topic information library; In response to the target network content matching any topic information in the topic information library and the topic color of the topic information matching the target network content being pure color, obtain the classification category of the topic information matching the target network content as the classification category of the target network content.

9. The method according to claim 3, characterized in that The obtaining of the multiple text samples includes: Obtain multiple text samples generated within a specified duration range before the current moment.

10. The method according to claim 1, characterized in that, The response that the target network content fails to match any topic information in the topic information library, and obtaining data samples based on the target network content includes: Obtain the classification category of the target network content obtained by classifying and annotating the target network content; Obtain the target network content and the classification category of the target network content as training samples; the training samples are used to update the parameters of the first classification model.

11. The method according to claim 10, wherein The method further includes: Process the target network content through the first classification model to obtain the first classification result of the target network content, and the first classification result is used to indicate the probability that the target network content belongs to various classification categories; Obtain a loss function value based on the difference between the first classification result and the classification category of the target network content; Update the parameters of the first classification model based on the loss function value.

12. A network content classification device, characterized in that, The device includes: A content acquisition module, configured to acquire target network content; A matching module, configured to perform a matching detection on the target network content and the topic information included in the topic information library; the topic information library includes the topic information corresponding to one or more topics; A sample acquisition module, configured to, in response to the target network content failing to match any topic information in the topic information library, acquire data samples based on the target network content; the data samples are used to update the topic information library and the first classification model; The first classification module is configured to, in response to the target network content matching any topic information in the topic information library, process the target network content through the first classification model, and obtain the classification category of the target network content output by the first classification model.

13. A computer device, characterized in that, The computer device includes a processor and a memory. At least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the network content classification method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the network content classification method according to any one of claims 1 to 11.

15. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the network content classification method according to any one of claims 1 to 11 is implemented.