Method, apparatus, device, and storage medium for processing publication information

HK40082728BActive Publication Date: 2026-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Patents
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2023-04-27
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, the efficiency and accuracy of content publishing feature analysis by content publishers are relatively low, making it difficult to handle a massive number of accounts.

Method used

By acquiring multiple historical publication information of the target object, the domain corresponding to each historical publication information is determined, and domain distribution information and content feature information are extracted and aggregated. The domain distribution information and content aggregation information are then used to determine the verticality of the target object.

Benefits of technology

It improves the accuracy and efficiency of content publisher feature analysis, enabling more accurate identification of explicit and implicit features of content publishers and enhancing the ability to determine the verticality of target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The application relates to the technical field of information processing, in particular to a method and device for processing published information, equipment and a storage medium, which can be applied to various scenes such as cloud technology, artificial intelligence, intelligent transportation and auxiliary driving. The method comprises the following steps: obtaining multiple historical published information of a target object, determining a field corresponding to each piece of historical published information; determining field distribution information corresponding to the target object; extracting features of each piece of historical published information to obtain content feature information corresponding to each piece of historical published information; performing aggregation processing on the multiple pieces of historical published information according to the content feature information corresponding to each piece of historical published information to obtain content aggregation information of the multiple pieces of historical published information; and determining the verticality of the target object based on the field distribution information and the content aggregation information. The application can improve the efficiency and accuracy of analyzing the content publishing features of a content publisher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, and in particular to a method, apparatus, device and storage medium for publishing information processing. Background Technology

[0002] In the new media landscape, platforms that allow users to voice their opinions, share, and disseminate information are called "self-media." The distribution method of displaying this media information through information feeds, and enabling user interaction with this information, has seen tremendous growth. High-quality content publishers and the high-quality content behind them have become the objects of competition among these platforms, because high-quality content attracts a large number of users and brings huge traffic. Content platforms offer various incentives to high-quality content publishers. Under this incentive, a large number of content reposting platforms have emerged. These reposting platforms modify, delete, and piece together original content published by other publishers, making it seem different from the original, but the most valuable core content remains the same.

[0003] In existing technologies, in order to analyze the content publishing characteristics of different content publishers, the characteristic information of the content publishers is usually determined manually by the operators based on the content publisher's reputation in the industry, its performance on other platforms, and the operators' personal experience. However, this method cannot handle a large number of accounts and has very low accuracy and processing efficiency. Summary of the Invention

[0004] The technical problem to be solved by this application is to provide a method, apparatus and electronic device for processing published information, which can improve the efficiency and accuracy of analyzing the content publishing characteristics of content publishers.

[0005] To address the aforementioned technical problems, this application provides a method for processing published information, the method comprising:

[0006] Obtain multiple historical postings of the target object and determine the domain corresponding to each historical posting;

[0007] Based on the domain corresponding to each historical published information, determine the domain distribution information corresponding to the target object;

[0008] Feature extraction is performed on each historical published information to obtain the content feature information corresponding to each historical published information;

[0009] Based on the content feature information corresponding to each historical published information, the multiple historical published information items are aggregated to obtain the content aggregate information of the multiple historical published information items.

[0010] Based on the domain distribution information and the content aggregation information, the verticality of the target object is determined; the verticality is used to characterize the degree to which the multiple historical published information items contain the same publishing characteristics.

[0011] On the other hand, this application provides a domain analysis apparatus for an object, the apparatus comprising:

[0012] The domain determination module is used to obtain multiple historical publication information of the target object and determine the domain corresponding to each piece of historical publication information.

[0013] The domain distribution information determination module is used to determine the domain distribution information corresponding to the target object based on the domain corresponding to each historical published information.

[0014] The content feature information generation module is used to extract features from each historical published information to obtain the content feature information corresponding to each historical published information;

[0015] The content feature information aggregation module is used to aggregate the multiple historical published information items according to the content feature information corresponding to each historical published information item, so as to obtain the content aggregation information of the multiple historical published information items;

[0016] The verticality determination module is used to determine the verticality of the target object based on the domain distribution information and the content aggregation information; the verticality is used to characterize the degree to which the multiple historical publication information contains the same publication characteristics.

[0017] On the other hand, this application provides a publishing information processing device, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the publishing information processing method as described above.

[0018] On the other hand, this application provides a computer storage medium storing at least one instruction or at least one program, wherein the at least one instruction or the at least one program is loaded by a processor and executed as described above in the information publishing processing method.

[0019] Implementing the embodiments of this application has the following beneficial effects:

[0020] This application determines the domain distribution information of a target object based on multiple historical publications. The domain distribution information can be determined based on the domain corresponding to each historical publication, thus representing the explicit characteristics of the target object. Content feature information corresponding to each historical publication is used to obtain the content aggregation information of each historical publication. The content feature information is the implicit information extracted from the historical publications, used to represent the implicit characteristics of the target object. By describing the target object using both explicit and implicit features, the accuracy of the target object's feature description can be improved. Therefore, by determining the verticality of the target object based on explicit and implicit features, rather than determining the domain characteristics of the target object manually, the accuracy and efficiency of determining the verticality of the target object can be improved. Attached Figure Description

[0021] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the implementation environment provided in the embodiments of this application;

[0023] Figure 2 This is a flowchart of a method for processing published information provided in an embodiment of this application;

[0024] Figure 3 This is a flowchart of a method for generating content feature information for video-related publishing information, provided in an embodiment of this application.

[0025] Figure 4 This is a flowchart illustrating a method for generating content feature information of text-based published information, as provided in an embodiment of this application.

[0026] Figure 5 This is a flowchart of a content aggregation information determination method provided in an embodiment of this application;

[0027] Figure 6 This is a flowchart of another method for determining content aggregation information provided in an embodiment of this application;

[0028] Figure 7 This is a flowchart of another content aggregation information determination method provided in the embodiments of this application;

[0029] Figure 8 This is a flowchart of a method for processing published information including audio information, provided in an embodiment of this application.

[0030] Figure 9 This is a flowchart of another method for processing published information including audio information provided in an embodiment of this application;

[0031] Figure 10 This is a flowchart of a target verticality fusion model generation method provided in an embodiment of this application;

[0032] Figure 11 This is a flowchart of a weight update method provided in an embodiment of this application;

[0033] Figure 12 This is a flowchart of an object management method provided in an embodiment of this application;

[0034] Figure 13 This is a schematic diagram of an information content publishing system provided in an embodiment of this application;

[0035] Figure 14 This is a schematic diagram of an information processing device provided in an embodiment of this application;

[0036] Figure 15 This is a schematic diagram of the electronic device structure provided in the embodiments of this application. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0039] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0040] As disclosed in this application, the information processing method or apparatus can be composed of multiple servers forming a blockchain, and the servers are nodes on the blockchain. Data during the information processing process can be stored on the blockchain.

[0041] First, the relevant terms used in the embodiments of this specification are explained as follows:

[0042] Articles: Articles recommended to users may include videos or pictures. These articles are usually edited and published by self-media after they start a public account.

[0043] Videos: Recommended videos for users to read, including vertical short videos and horizontal short videos, provided in the form of a feed stream.

[0044] PGC (Professional Generated Content): An internet term referring to professionally produced content (video websites) and expert-produced content (microblogging). It broadly refers to personalized content, diverse perspectives, democratized dissemination, and virtualized social relationships. Also known as PPC (Professionally-produced Content).

[0045] MCN (Multi-Channel Network): This is a product form of multi-channel network that combines PGC content and, with strong capital support, ensures the continuous output of content, thereby ultimately achieving stable commercial monetization.

[0046] UGC (User Generated Content) refers to user-generated content, which emerged alongside the Web 2.0 concept, which emphasizes personalization. It's not a specific business service, but rather a new way users engage with the internet, shifting from primarily downloading to a balance between downloading and uploading.

[0047] PUGC (Professional User Generated Content): This refers to professional audio content produced in UGC format that is relatively close to PGC.

[0048] Feeds: A news feed, also translated as source, delivery, information provision, summary, source, news subscription, or web feed (English: web feed, news feed, syndicated feed), is a data format that websites use to disseminate the latest information to users. It is usually arranged in a timeline format; timeline is the most original, intuitive, and basic form of feed display. A prerequisite for users to subscribe to a website is that the website provides news sources. The aggregation of feeds in one place is called aggregation, and the software used for aggregation is called an aggregator. For end users, an aggregator is software specifically designed for subscribing to websites, and is also commonly referred to as an RSS reader, feed reader, or news reader.

[0049] Machine Learning (ML) is a multidisciplinary field that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.

[0050] Please see Figure 1 The illustration shows an implementation environment provided in this application embodiment. The implementation environment may include: a content uploading terminal 110, a content publishing platform 120, and a content display terminal 130. The content uploading terminal 110, the content publishing platform 120, and the content display terminal 130 can communicate with each other via a network.

[0051] Specifically, each user can create content to be published through the content upload terminal 110 and upload it to the corresponding content publishing platform 120. The content publishing platform 120 can review, classify, and distribute the content information published by the content upload terminal 110. Specifically, when distributing information, the content can be distributed to the content display terminal 130. The content upload terminal 110 and the content display terminal 130 can be the same terminal or different terminals.

[0052] The content uploading terminal 110 or the content display terminal 130 can communicate with the content publishing platform 120 based on a browser / server (B / S) mode or a client / server (C / S) mode. The content uploading terminal 110 or the content display terminal 130 may include physical devices such as smartphones, tablets, laptops, digital assistants, smart wearable devices, in-vehicle terminals, and servers, and may also include software running on the physical devices, such as applications. The operating system running on the content uploading terminal 110 or the content display terminal 130 in this embodiment may include, but is not limited to, Android, iOS, Linux, and Windows.

[0053] The content publishing platform 120 can establish a communication connection with the content uploading terminal 110 or the content display terminal 130 via wired or wireless means. The content publishing platform 120 may include a stand-alone server, a distributed server, or a server cluster consisting of multiple servers, wherein the server may be a cloud server.

[0054] To address the issues of low accuracy and processing efficiency in determining the feature information of content objects in existing technologies, this application provides a method for processing published information. Please refer to... Figure 2 The method may include:

[0055] S210. Obtain multiple historical publication information items of the target object and determine the domain corresponding to each historical publication information item.

[0056] The object in this application embodiment can be an information object or a content object, specifically including the account that publishes content, i.e., the account holder registered on different content publishing platforms. Historical publishing information refers to content information already published by the target object. In this application embodiment, to facilitate the analysis of the target object and reduce the amount of historical publishing information to analyze, historical publishing information of the target object within a certain period can be selected for analysis. Specifically, this can be done by: obtaining a set of historical publishing information of the target object within a preset time period adjacent to the current time; and randomly selecting a preset number of historical publishing information from the set of historical publishing information. Since content information closer to the current time better reflects the current publishing characteristics of the target object, selecting content information closer to the current time can also improve the accuracy of the publishing characteristic analysis of the target object.

[0057] For content publishing platforms, a corresponding content category system can be set up. For example, categories could include primary categories such as TV dramas, variety shows, entertainment, society, current events, news, internet celebrities, sports, technology, games, and animation. After content is uploaded, during the main content processing flow, the content is typically categorized using both machine and manual methods. In practice, primary category classification information can be used for labeling. By analyzing the category labeling information of historical posts, the publishing category to which each historical post belongs can be determined. The publishing category to which the content information belongs can be used to characterize the explicit features of the object.

[0058] S220. Based on the domain corresponding to each historical published information, determine the domain distribution information corresponding to the target object.

[0059] Domain distribution information can be used to characterize the degree of concentration of the domains corresponding to the historical published information of a target object; if the number of domains corresponding to the historical published information of a target object is small, it indicates that the domain concentration of the target object is high; if the number of domains corresponding to the historical published information of a target object is large, it indicates that the domain concentration of the target object is low.

[0060] Specifically, the method for determining domain distribution information may include: determining the number of historical publications corresponding to each category identifier based on the category identifier corresponding to each historical publication; and determining the domain distribution information corresponding to the target object based on the number of historical publications corresponding to each category identifier. Here, the domain corresponding to each historical publication is determined by the category identifier corresponding to each historical publication, with one category identifier corresponding to one publication domain category. Therefore, the domain to which the historical publication belongs can be determined based on the category identifier corresponding to the historical publication.

[0061] Furthermore, the domain distribution information can be specifically represented in the form of information entropy, which may include: calculating the category information entropy corresponding to the target object based on the number of historical published information corresponding to each category identifier; and determining the domain distribution information corresponding to the target object based on the category information entropy. Specifically, the ratio of published information corresponding to each published domain category is obtained by comparing the number of historical published information under each published domain category with the total number of selected historical published information. Based on this ratio, the category information entropy corresponding to the target object is calculated. The specific calculation result of the category information entropy is shown in equation (1):

[0062]

[0063] Where i represents the i-th publishing domain category, n represents the total number of domain categories in multiple historical publishing information, and Pi represents the ratio of the above to the publishing information corresponding to each publishing domain.

[0064] Category information entropy tends to identify objects whose publishing domains are relatively concentrated as having high creative focus, thus indicating high verticality. Category information entropy is determined based on the explicit characteristics of objects; that is, it can be determined by the domain classification of the published content.

[0065] If the category information entropy is large, it indicates that the category concentration is low; if the category information entropy is small, it indicates that the category concentration is high. Therefore, the category concentration information corresponding to the target object can be determined as 1 - category information entropy.

[0066] S230. Extract features from each historical published information to obtain content feature information corresponding to each historical published information.

[0067] Content feature information can be used to characterize the implicit features in the content information published by an object. In the specific implementation process, deep learning algorithms can be used to generate content feature information, and different content feature information generation methods can be used for different forms of content information.

[0068] Please see Figure 3 It illustrates a method for generating content feature information for video-related publishing information, which may include:

[0069] S310. When the historical published information is video information, the video information is divided into video frames to obtain a video frame sequence corresponding to the video information.

[0070] S320. Extract multiple key video frames from the video frame sequence and generate image feature vectors corresponding to each key video frame.

[0071] S330. Perform feature fusion on the image feature vectors corresponding to each key video frame to obtain content feature information corresponding to the video information.

[0072] Video-related published information can include multiple video frames. The published information can be divided into video frames to obtain a video frame sequence. Key video frames can be determined by calculating the similarity between adjacent video frames. Specifically, this involves calculating the similarity between adjacent video frames in the sequence; if the similarity between a subsequent video frame and a preceding video frame is less than a preset similarity, the subsequent video frame is identified as a key video frame. By identifying and processing key video frames, the number of video frames to be processed can be reduced while preserving the feature information in the video-related published information, avoiding the loss of feature information, thereby improving the efficiency of generating content feature information for video-related features.

[0073] However, many content-specific vertical characteristics cannot be reflected by explicit classification features. In such cases, implicit features can be used. Specifically, for example, the same scene, the same person in the same video, or the same visual style, implicit features of content information are needed. Implicit features can include image information features, text information features, audio information features, etc. Content-based "implicit" features: using video content embedding. This content embedding generally has two layers. The first layer means representation learning, low-dimensional dense features. The second layer means metric learning, similarity measurement vectors. The "distance" between two vectors represents the "similarity" between two objects. The content embedding generation process includes: inputting a video frame sequence, extracting keyframes through a TSN (Temporal Segment Network), then extracting image features through Xception, and finally using the image feature vectors obtained from the intermediate layers of the multimodal video classification model NeXtVLad. The image feature vectors are then summed and averaged to obtain the video vector embedding.

[0074] Please see Figure 4 It illustrates a method for generating content feature information of text-based published information, which may include:

[0075] S410. When the historical published information is text information, the text information is processed to remove text styles.

[0076] S420. Extract semantic features from the text information whose text style has been removed to obtain content feature information corresponding to the text information.

[0077] For text-based published information, the focus is on the specific text content, thus requiring deformatting to obtain the main text information. In this embodiment, the BERT (Bidirectional Encoder Representations from Transformers) model can be used to generate content feature information corresponding to the text information. For example, for image-text published information, formatting style text such as various HTML tags is removed from the main text. The semantic features of the main text information are then extracted using the BERT model, that is, the text string is converted into a vector, and the vector from the penultimate layer of the BERT model is extracted as the content feature vector corresponding to the text information. Therefore, by deformatting text-based published information, corresponding content feature information is generated based on the formatted text information, thereby further improving the efficiency and accuracy of generating content feature information for text-based published information.

[0078] S240. Aggregate the multiple historical releases based on the content feature information corresponding to each historical release to obtain the content aggregation information of the multiple historical releases; wherein the aggregation process may specifically be to calculate the similarity of the content feature information corresponding to each historical release.

[0079] In this embodiment, the content aggregation information can be used to characterize the degree of aggregation among the various historical posts published by the target object; the higher the degree of aggregation, the higher the similarity of the various historical posts; conversely, the lower the degree of aggregation, the lower the similarity of the various historical posts.

[0080] Please see Figure 5 It illustrates a method for determining content aggregation information, which may include:

[0081] S510. Combine the content feature information corresponding to each of the historical published information items in pairs to obtain multiple content feature information groups.

[0082] S520. Calculate the similarity between two content feature information in each content feature information group.

[0083] S530. Based on the similarity to each of the content feature information groups, obtain the content aggregation information of each historical publication information.

[0084] The content feature information identified above can specifically be content feature vectors, i.e., vectors used to characterize the content features of published information. The similarity between the content feature vectors of any two historical published information items is calculated. Here, the cosine similarity can be calculated using the cosine formula to calculate the distance between two vectors, and the distance between the two vectors is used as the similarity between them. The larger the distance between two vectors, the smaller the similarity; the smaller the distance between two vectors, the larger the similarity. Therefore, the distance between vectors and the similarity between vectors are negatively correlated. The reciprocal of the distance between two vectors can be used to determine the similarity between the two vectors, or 1 minus the distance between two vectors can be used to determine the similarity between the two vectors. Correspondingly, the sum of the similarities corresponding to each content feature information group can be used to determine the content aggregation information, i.e., the content aggregation information is:

[0085]

[0086] or

[0087] s2=∑(1-distance between the two vectors) (3)

[0088] Please see Figure 6 It illustrates another method for determining content aggregation information, which may include:

[0089] S610. The content feature information corresponding to each historical published information is averaged to obtain the central content feature information of each content feature information.

[0090] S620. Calculate the similarity between the content feature information corresponding to each historical published information item and the central content feature information.

[0091] S630. Based on the similarity between the content feature information corresponding to each historical published information and the central content feature information, the content aggregation information of each historical published information is obtained.

[0092] In this embodiment, the central content feature information can also be determined based on the content feature vectors corresponding to each historical release information. This can be achieved by summing the content feature vectors corresponding to each historical release information and then taking the average. The similarity between the content feature information corresponding to each historical release information and the central content feature information can be calculated. The similarity calculation can refer to the vector similarity calculation method described above in this embodiment. Correspondingly, the content aggregation information can also be the sum of the similarity calculation results.

[0093] Please see Figure 7 It illustrates yet another method for determining content aggregation information, which may include:

[0094] S710. Cluster the content feature information corresponding to each of the historical release information to obtain multiple cluster center vectors.

[0095] S720. Combine the plurality of cluster center vectors in pairs to obtain at least one cluster vector group.

[0096] S730. Calculate the similarity between two cluster center vectors in the cluster vector group.

[0097] S740. Based on the similarity to the clustering vector group, obtain the content aggregation information of each historical publication information item.

[0098] In this embodiment, KNN (K-Nearest Neighbor) can also be used to cluster the content feature vectors corresponding to each historical published information to obtain k cluster center vectors; then the similarity calculation method described above is used to calculate the similarity between any two cluster center vectors, and the corresponding content aggregation information can also be the sum of the similarity calculation results.

[0099] In the specific implementation of the methods for determining content aggregation information listed in this embodiment, any one of the methods can be selected, or two or more methods can be selected and combined. When two or more methods are used to determine content aggregation information, the results of the content aggregation information of the two or more methods can be averaged. This embodiment does not make specific limitations.

[0100] Specifically, at least two of the historical published information items include audio information; these at least two historical published information items may also include image information and text information; content feature information analysis can be performed based on the audio information of these at least two historical published information items; please refer to [link / reference] for details. Figure 8 It illustrates a method for processing published information containing audio information, which may include:

[0101] S810. Perform acoustic feature extraction on the audio information of the at least two historical release information items to obtain acoustic feature information corresponding to the at least two historical release information items.

[0102] S820. Based on the acoustic feature information, obtain the content feature information corresponding to the at least two historical publication information items.

[0103] S830. Aggregate the content feature information corresponding to the at least two historical published information items to obtain the content aggregate information of the at least two historical published information items.

[0104] The content aggregation information is used to characterize the degree to which the at least two historical published information items contain the same acoustic feature information.

[0105] Among them, acoustic feature information is the physical quantity of the acoustic properties of speech, and it is also a general term for the acoustic performance of various sound elements. Such as the energy concentration area, formant frequency, formant intensity and bandwidth representing timbre, and the duration, fundamental frequency and average speech power representing the prosodic characteristics of speech. Specifically, by using two acoustic feature information, it can be determined whether the sound signals corresponding to these two acoustic features are produced by the same person, or whether they are the same instrument, or whether they are the same or similar music, etc.

[0106] By aggregating the acoustic feature information corresponding to the historical published information, it is possible to determine the degree to which these at least two historical published information contain the same acoustic feature information. The higher the degree of content aggregation, the higher the similarity of the acoustic feature information contained in these at least two historical published information; conversely, the lower the degree of aggregation, the lower the similarity of the acoustic feature information contained in these at least two historical published information.

[0107] Please see Figure 9 It illustrates another method for processing published information that includes audio information, which may include:

[0108] S910. Perform speech recognition on the audio information of the at least two historical release information items to obtain speech recognition information corresponding to the at least two historical release information items.

[0109] S920. Based on the speech recognition information, obtain content feature information corresponding to the at least two historical publication information items.

[0110] S930. Aggregate the content feature information corresponding to the at least two historical published information items to obtain the content aggregate information of the at least two historical published information items.

[0111] The content aggregation information is used to characterize the degree to which the at least two historical published information items contain the same speech recognition information.

[0112] When the audio information is specifically language-based information such as dialogue or narration, the speech recognition information can include the text content information in the audio information. Therefore, by processing the speech recognition information corresponding to the historical release information, it is possible to determine the degree to which these at least two historical release information contain the same speech recognition information. The higher the degree of content aggregation, the higher the similarity of the speech recognition information contained in these at least two historical release information; conversely, the lower the degree of aggregation, the lower the similarity of the speech recognition information contained in these at least two historical release information.

[0113] Specifically, based on the aforementioned acoustic feature information, corresponding acoustic feature vectors can be generated; based on the aforementioned speech recognition information, corresponding speech feature vectors can be generated. This embodiment can be used when performing content aggregation based on the acoustic feature vectors corresponding to historically published information, or when performing content aggregation based on the speech feature vectors corresponding to historically published information. Figures 5-7 Any method implementation is acceptable, and no specific method is required here.

[0114] Since historical published information may include image information, audio information, and text information, the content feature information of the target object may have multiple dimensions, namely the content feature information corresponding to image information, the content feature information corresponding to audio information, and the content feature information corresponding to text information. Therefore, when performing content aggregation, content feature information of the same type can be aggregated. That is, the content aggregation information corresponding to multiple historical published information also has multiple dimensions, namely the content aggregation information corresponding to image information, the content aggregation information corresponding to audio information, and the content aggregation information corresponding to text information.

[0115] Specifically, since the content information published by the same object is not of a single type, but may include text and image types, short video types, etc., verticality can be calculated for each track. Specifically, the content information feature model mentioned above can be used to calculate the content aggregation information corresponding to text and image types, short video types, and short video types, etc. When finally determining the verticality, the content aggregation information corresponding to each type can be weighted and summed. The specific weights can be determined based on the content information corresponding to each type. For example, if the number of content information corresponding to text and image types, short video types, and short video types are a1, a2, and a3, respectively, the corresponding weights are a1 / (a1+a2+a3), a2 / (a1+a2+a3), and a3 / (a1+a2+a3). The more types there are, the greater the weight, thus making their influence on the final calculated verticality.

[0116] S250. Based on the domain distribution information and the content aggregation information, determine the verticality of the target object; the verticality is used to characterize the degree to which the multiple historical published information items contain the same publishing characteristics.

[0117] Shared publishing characteristics refer to the consensus features of various historical publishing messages. These can include belonging to the same category (domain), the same scenario, the same theme, the same style, the same voice narration, or similar background music. When any of these consensus features are present in the historical publishing messages of a target object, the target object can be determined to be vertical; verticality can be used to measure the degree to which multiple publishing messages contain the same consensus feature. For example, corresponding to Figure 8 and Figure 9 The same release features can be the same acoustic feature information or the same speech recognition information.

[0118] The domain distribution information in this application embodiment can be used to characterize the dispersion of domains marked by historically published information. The higher the category information entropy corresponding to the published domain, the more dispersed the domains marked by historically published information are; conversely, the more concentrated the marked domains are. Verticality reflects the degree to which an object focuses on a domain. The higher the verticality, the more concentrated the creative domain is; conversely, the more dispersed the creative domain is. To ensure that the meaning of category information entropy and domain verticality are consistent, the category information entropy can be transformed when determining verticality. For example, the transformation result can be s1 = 1 - category information entropy. After determining the domain distribution information and content aggregation information of the target object, the verticality corresponding to the target object can be obtained through the target verticality fusion model. Please refer to [link to relevant documentation]. Figure 10 It illustrates a method for generating a target verticality fusion model, which may include:

[0119] S1010. Obtain the verticality level of multiple annotation objects.

[0120] S1020. Using a preset verticality fusion model, verticality prediction is performed based on the domain distribution information and content aggregation information of the multiple labeled objects to obtain the predicted verticality of the multiple labeled objects; the preset verticality fusion model includes a first weight corresponding to the domain distribution information and a second weight corresponding to the content aggregation information.

[0121] S1030. Based on the verticality levels of the multiple labeled objects and the predicted verticality of the multiple labeled objects, update the first weight and the second weight to obtain the target verticality fusion model.

[0122] The preset verticality fusion model can be:

[0123] s=α*s1+(1-α)*s2 (4)

[0124] For an explanation of the relevant parameters in equation (4), please refer to the above content of this embodiment. α*s1 corresponds to the domain distribution information model, (1-α)*s2 corresponds to the content information feature model, α corresponds to the first weight, and (1-α) corresponds to the second weight.

[0125] The first and second weights in the preset verticality fusion model are updated by using the verticality levels of multiple labeled objects and the predicted verticality, so that the preset verticality model is updated in the direction of the labeled verticality.

[0126] Further, please refer to Figure 11 It illustrates a weight update method, which may include:

[0127] S1110. Based on the verticality level of the multiple labeled objects, determine the number of labeled objects at each verticality level.

[0128] S1120. Sort the predicted verticality of the multiple labeled objects to obtain the verticality sorting result.

[0129] S1130. Based on the number of labeled objects at each verticality level, determine the predicted objects corresponding to each verticality level from the verticality ranking results.

[0130] S1140. Based on the labeled objects corresponding to each verticality level and the predicted objects corresponding to each verticality level, determine the precision and recall of the preset verticality fusion model.

[0131] S1150. Update the first weight and the second weight based on the precision and recall rate to obtain the target verticality fusion model.

[0132] Precision and recall can include accuracy and recall, and can be used to measure the predictive performance of a model. For specific calculation methods of the model's accuracy and recall, please refer to the methods in the prior art, which will not be repeated in this embodiment. When the precision and recall of the preset verticality fusion model meet the preset conditions, it can be determined as the target verticality fusion model. In specific implementation, α in the target verticality fusion model can be taken as 0.2.

[0133] Furthermore, a target verticality fusion model can be used to predict the verticality of multiple labeled objects to obtain the target verticality of multiple labeled objects; based on the target verticality of the labeled objects at each verticality level, the threshold for dividing verticality levels can be determined.

[0134] Specifically, a validation set can be manually labeled. A batch of objects are labeled in advance by referring to the verticality definition standard as the validation set. When the fusion model uses the validation set, the corresponding level classification threshold can be determined, and the precision and recall of (non-vertical / vertical / general accounts) can be given respectively, which serve as the benchmark for the model's ability to calculate account verticality.

[0135] Verticality definition annotation:

[0136] (1) If, within 3 months, for example, 10 videos are randomly selected (text and image content is handled in a similar way), more than 90% of the videos meet any of the following definitions of "vertical", then the object is considered vertical.

[0137] (2) The labeled fields are: TV series, variety shows, entertainment, society, current affairs, news, internet celebrities, sports, technology, games, animation and other commonly used information flow platform classification systems.

[0138] (a) Vertical sound: Definition: Commentary class, all content of the object comes from the same person's commentary.

[0139] (b) Scene verticality: Definition: The videos are all shot in the same scene, such as all in a mountain or all in a dormitory.

[0140] (c) Vertical presentation style: Definition: Video content can be very diverse, but they all use the same presentation techniques, such as using the same anime characters, or using a certain fixed format, such as using maps, etc.

[0141] (d) Verticality of a certain important attribute: Definition: Very vertical in a certain aspect, such as style, tone, etc.

[0142] (e) Vertical Theme: Definition: The videos published can be categorized as being based on a certain theme. The definition of a theme is very broad and can be certain things, certain objects, etc.

[0143] In some implementations, multiple target objects can be classified into three categories by manual verticality annotation, with verticality from high to low: vertical, general, and non-vertical. Verticality "vertical" can be understood as: the object can be clearly determined to be vertical based on the above definition of verticality annotation; general verticality can be understood as: the object cannot be clearly determined to be vertical based on the above definition of verticality annotation, it may be vertical or it may not be vertical; and non-vertical verticality can be understood as: the object can be clearly determined to be non-vertical based on the above definition of verticality annotation.

[0144] For example, based on the above definition of verticality, a test object set can be constructed, which may include N objects, specifically accounts. There are 'a' vertical objects, 'b' general objects, and 'c' non-vertical objects, where a + b + c = N. Here, "general objects" refers to objects that are difficult to determine. The main purpose here is to automatically verify the model's characterization results. The model's output results can also be manually checked one by one.

[0145] During the update process of the preset verticality fusion model, the preset verticality fusion model is used to predict the verticality of N objects in the test object set, obtaining the predicted verticality of these N objects. The N objects are then sorted in descending order of predicted verticality. The top 'a' objects are identified as predicted vertical objects, the middle 'b' objects are identified as predicted general objects, and the bottom 'c' objects are identified as predicted non-vertical objects. The labeled 'a' vertical objects are matched with the 'a' predicted vertical objects, the labeled 'b' general objects are matched with the 'b' predicted general objects, and the labeled 'c' non-vertical objects are matched with the 'c' predicted non-vertical objects, yielding the precision and recall of the current preset verticality fusion model. If the precision and recall do not meet the preset conditions, the first and second weights are adjusted, and the steps of verticality prediction and precision and recall calculation are repeated until the verticality fusion model meets the preset conditions.

[0146] For each labeled object, the target verticality fusion model described above outputs a corresponding verticality score. These scores are sorted from highest to lowest. Then, based on the verticality level of the labeled object, a threshold for each verticality level is determined. Specifically, the a-th and a+1-th sorted objects are the dividing points between vertical and general objects, while the a+b-th and a+b+1-th sorted objects are the dividing points between general and non-vertical objects. A score greater than a first preset value (thres1) indicates a vertical domain, and a score less than a second preset value (thres2) indicates a non-vertical domain. The first preset value (thres1) is greater than the second preset value (thres2). The first preset value (thres1) can be the average of the verticality scores of the a-th and a+1-th sorted objects, and the second preset value (thres2) can be the average of the verticality scores of the a+b-th and a+b+1-th sorted objects.

[0147] Please see Figure 12 It illustrates an object management method that may include:

[0148] S1210. Determine the verticality level of the target object based on the verticality of the target object and the verticality level classification threshold.

[0149] S1220. Based on the verticality level of the target object, control the publication information of the target object.

[0150] Once the verticality level threshold is determined, a target verticality fusion model can be used to determine the verticality of any object, thus obtaining the object's verticality level. This allows for a more granular grading mechanism for objects, enabling increased subsidies and incentives for objects with high verticality, guiding authors to focus more on providing high-quality content, while reducing subsidies and incentives for objects with lower verticality. Simultaneously, in the content review process, due to limited review resources, objects with low verticality are placed at the end of the review scheduling process, optimizing review scheduling and improving review efficiency. This effectively controls low-verticality content reposting, reduces the impact on the content publishing platform's development, and promotes the healthy development of the object ecosystem.

[0151] This application integrates explicit and implicit features to describe the verticality of an object. Explicit features are determined based on the classification domains labeled by the object's historical publishing information, which is simple to calculate and easy to interpret. Implicit features are based on the embedding vectors of the object's historical publishing content. They can measure verticality from dimensions other than domain classification. Some objects publish content that is not often classified into the same category, but this content information has some commonalities. The content embedding vector here essentially measures the similarity of the content. Some objects may publish content that is domain-vertical but not very similar. For example, an anime object may publish different anime works with low similarity, but the domain is consistent, all being anime. Overall, it is a domain-vertical object. Therefore, the verticality characterization from the perspectives of domain information entropy and content aggregation information can be complementary. Correspondingly, there are domain information entropy models and content feature information models. The fusion model after combining these two models can achieve greater performance improvement in determining the verticality of an object.

[0152] This application determines the domain distribution information of a target object based on multiple historical publications. The domain distribution information can be determined based on the domain corresponding to each historical publication, thus representing the explicit characteristics of the target object. Content feature information corresponding to each historical publication is used to obtain the content aggregation information of each historical publication. The content feature information is the implicit information extracted from the historical publications, used to represent the implicit characteristics of the target object. By describing the target object using both explicit and implicit features, the accuracy of the target object's feature description can be improved. Therefore, by determining the verticality of the target object based on explicit and implicit features, rather than determining the domain characteristics of the target object manually, the accuracy and efficiency of determining the verticality of the target object can be improved.

[0153] Please see Figure 13 It illustrates an information content publishing system, and the functions of each service module of the system are as follows:

[0154] 1. Content production end and content consumption end

[0155] (1) PGC or UGC, MCN or PUGC content producers provide text and image content or upload video content, including short videos and mini videos, through mobile terminal or backend interface API system. These are the main sources of content for distribution.

[0156] (2) By communicating with the upstream and downstream content interface servers, first obtain the upload server interface address, and then publish the content;

[0157] (3) As a consumer, communicate with the upstream and downstream content interface servers to obtain the index information of the accessed content, and then communicate with the upstream and downstream content interface servers and the content exit service to directly consume the content. The prerequisite for consumption is to obtain the index of the content through Feeds recommendation distribution, i.e. the entry address for content access.

[0158] (4) Feeds and user click behavior and environment reporting module: collect the user's current network environment and the user's click operation behavior on Feeds intermediate information and the exposure data of Feeds content, and report them to the statistics reporting interface server.

[0159] (5) If the video content is reported, the video playback time, cache time and various interactive behaviors of the content such as comments, forwarding, sharing, collection, likes, reports, etc. are reported;

[0160] 2. Uplink and downlink content interface servers and content export services

[0161] (1) Communicate directly with the content production end. The content submitted from the front end usually includes the title, publisher, summary, cover image, and publication time. Store the content in the database.

[0162] (2) The content delivery service communicates with the recommendation and distribution system to obtain the recommendation and distribution results and send them to the consumer end to be displayed in the user's Feeds list;

[0163] (3) Content export services are usually a set of access services that are geographically located near the user;

[0164] (4) At the same time, report the posting information of each account to the statistics interface server, including the posting time and content type. Also, store the content tagging information provided by the self-media, such as category, tags, selected cover image, and title, as extended information in the content database.

[0165] 3. Content database and content storage service

[0166] (1) The core database of content, where all the metadata of the content published by producers is stored. The focus is on the metadata of the content itself, such as size, cover image link, title, publication time, account author, source channel, entry time, and also the classification of content during the manual review process (including first, second and third level classification and tag information, such as an article explaining ×× mobile phone, the first level classification is technology, the second level classification is smartphone, the third level classification is domestic mobile phone, and the tag information is ××).

[0167] (2) During the manual review process, information in the content database will be read, and the results and status of the manual review will also be sent back to the content database for storage.

[0168] (3) Content processing in the entire business process mainly includes machine processing and manual review. Based on different content tags, the content library is divided into different content pools. Recommendation distribution server and deduplication server, as well as content feature modeling service, all need to obtain content from the content database. For example, the image and text deduplication server will load content that has been entered and used in the past period (such as a week) according to business needs. For duplicate content that is re-entered into the database, a filter tag will be added and it will no longer be provided to the content recommendation service for output to users.

[0169] (4) Deduplication service is a machine processing process, and the results are stored in the content database;

[0170] (5) Content storage service stores the actual source files of the content, such as the source files of videos, the original files of images, etc., to provide the original information for building the embedding of the content;

[0171] 4. Dispatch Center

[0172] (1) Responsible for the entire scheduling process of content flow, receiving the content into the database through the upstream and downstream content interface servers, and then obtaining the content's metadata from the content database;

[0173] (2) Schedule the deduplication server to mark and filter duplicate entries;

[0174] (3) For content that cannot be processed by machines, such as politically sensitive or security issues that require manual review, call the manual review system for manual review.

[0175] 5. Manual review service system

[0176] (1) It is necessary to read the original information of the video content itself in the content database. This is usually a complex web database-based system. The main purpose is to ensure that the pushed content complies with the access permitted by local laws and policies, such as whether it is sensitive information and perform a preliminary filtering.

[0177] (2) The content reviewed comes from self-media accounts' proactive posts and supplementary information obtained from public networks by web crawlers;

[0178] (3) The review results are finally written into the content database through the dispatch center service;

[0179] (4) Assist the machine processing process by classifying, tagging and correcting the content based on the machine processing, and finally storing the results in the content database;

[0180] 6. De-duplication service

[0181] (1) Communication with the content scheduling server mainly includes title deduplication, cover image deduplication, content text deduplication, and video and audio fingerprint deduplication. Usually, the title and text of the image and text content are vectorized, and simmhash and BERT text vectors are used for image vector deduplication. For video content, video fingerprints and audio fingerprints are extracted to construct vectors, and then the distance between vectors, such as Euclidean distance, is calculated to determine whether there is a repetition.

[0182] 7. Statistical reporting interface server

[0183] (1) Receive reports from content consumer users on their current network environment, click behavior of users on Feeds intermediate information, and exposure data of Feeds articles;

[0184] (2) Write the reported statistical data results into the statistical database;

[0185] (3) Accept the original transaction data of accounts published by the content production portal;

[0186] 8. Verticality Feature Fusion Model

[0187] (1) Following the specific fusion method described above, the results of the information entropy verticality model and the verticality characterization model based on account content embedding are fused together to form the final model for characterizing the verticality of self-media accounts.

[0188] 9. Account Verticality Recognition Service

[0189] (1) Implement the above-mentioned verticality feature fusion model in an engineering manner;

[0190] (2) Communicate with the dispatch center service to complete the vertical identification of the issuing account;

[0191] 10. Account Content Vector Service

[0192] (1) Construct content vector embeddings for the text and video content published by the account, which serve as the input dimension for measuring the verticality of the account. For the specific methods and models for constructing vectors, please refer to the description above.

[0193] (2) Communicate with the account verticality fusion model to provide results that implicitly measure account verticality;

[0194] 11. Statistical Database

[0195] (1) Receive statistical data reports from content consumption terminals to provide data support for subsequent statistical analysis and mining;

[0196] (2) Receive the document production log report from the content production end.

[0197] The object domain analysis method and system provided in this application can be used to characterize the degree to which the content published by an object belongs to the same category or the same theme. This serves as a basic feature of the object, a necessary condition for object authentication, and a fundamental input feature for object quality modeling such as object grading and object plagiarism identification, thus improving the overall account ecosystem. Furthermore, it introduces verticality for different content formats, which can be used for tiered grading. The same object may publish text and images, short videos, and other content; a single grading level might not be sufficient. Therefore, an account is evaluated for tiers across different tiers: text and images, short videos, and other short videos. The overall grading level is a combination of the results from all tiers. The core idea is to model the explicit and implicit features of the object's published content separately, then fuse these two types of features into a model. This fused model characterizes the object's verticality. Explicit features are primarily based on the information entropy of the vertical category information of the object's recently published content, while implicit features mainly use the account's published content and embedding vectors constructed from the content itself. Vector clustering is then used to characterize the account's verticality. Finally, a manually labeled account verticality verification dataset is used to determine a suitable model threshold as the final parameter. This application allows for direct application of the necessary conditions for certification of self-media account owners in content production fields (games / anime / film and television, etc.), requiring a certain degree of verticality. Furthermore, since verticality reflects the level of focus of an object, during recommendation and distribution, objects with low verticality will be given lower priority, restricted, or even canceled, concentrating traffic on content creators who are truly dedicated. Verticality can provide a more granular grading and characterization mechanism based on content scenarios. Objects with low verticality will be placed at the end of the review and scheduling process, optimizing review and scheduling and improving the account ecosystem. Finally, as an auxiliary feature, it can better characterize the level and quality features of an object, leveraging the role of basic features.

[0198] Typical application scenarios of account verticality include: (1) Direct application: The necessary condition for the certification of self-media account owners in the content production field (game / anime field) is that the verticality must be greater than a certain threshold to reflect the author's level of focus.

[0199] (2) Indirect application: as input features for other account models, such as account level models.

[0200] (3) Verticality by domain, for example, for classifying by track. For self-media accounts, the same account may publish text, images, short videos and short videos. If a single level is not enough to describe it, the level of an account is evaluated by track: text track, short video track and short video track. The verticality of each track can be described separately by different domain results.

[0201] (4) Verticality can reflect an author’s focus on content creation and can be used in conjunction with product operation strategies, such as demotion of recommendation distribution and demotion of review scheduling in the content processing chain.

[0202] This embodiment also provides an information publishing processing device; please refer to [link / reference]. Figure 14 The device may include:

[0203] Domain determination module 1410 is used to obtain multiple historical publication information of the target object and determine the domain corresponding to each piece of historical publication information.

[0204] The domain distribution information determination module 1420 is used to determine the domain distribution information corresponding to the target object based on the domain corresponding to each historical published information.

[0205] The content feature information generation module 1430 is used to extract features from each historical published information to obtain the content feature information corresponding to each historical published information.

[0206] The content feature information aggregation module 1440 is used to aggregate the multiple historical published information according to the content feature information corresponding to each historical published information to obtain the content aggregation information of the multiple historical published information.

[0207] Verticality determination module 1450 is used to determine the verticality of the target object based on the domain distribution information and the content aggregation information; the verticality is used to characterize the degree to which the multiple historical published information contains the same publishing characteristics.

[0208] Furthermore, the domain distribution information determination module 1420 includes:

[0209] The quantity determination module is used to determine the quantity of historical published information corresponding to each category identifier based on the category identifier corresponding to each historical published information; wherein, the domain corresponding to each historical published information is determined by the category identifier corresponding to each historical published information;

[0210] The first determining module is used to determine the domain distribution information corresponding to the target object based on the number of historical published information corresponding to each category identifier.

[0211] Furthermore, the first determining module includes:

[0212] The category information entropy calculation module is used to calculate the category information entropy corresponding to the target object based on the number of historical published information corresponding to each category identifier;

[0213] The second determining module is used to determine the domain distribution information corresponding to the target object based on the category information entropy.

[0214] Furthermore, at least two of the historical published information items include audio information;

[0215] The content feature information generation module 1430 includes:

[0216] An acoustic feature extraction module is used to extract acoustic features from the audio information of the at least two historical release information items to obtain acoustic feature information corresponding to the at least two historical release information items;

[0217] The first generation module is used to obtain content feature information corresponding to the at least two historical publication information based on the acoustic feature information;

[0218] The content feature information aggregation module 1440 includes:

[0219] The first aggregation module is used to aggregate the content feature information corresponding to the at least two historical published information items to obtain the content aggregation information of the at least two historical published information items.

[0220] The content aggregation information is used to characterize the degree to which the at least two historical published information items contain the same acoustic feature information.

[0221] Furthermore, at least two of the historical published information items include audio information;

[0222] The content feature information generation module 1430 includes:

[0223] The speech recognition module is used to perform speech recognition on the audio information of the at least two historical release information items to obtain speech recognition information corresponding to the at least two historical release information items;

[0224] The second generation module is used to obtain content feature information corresponding to the at least two historical publication information based on the speech recognition information;

[0225] The content feature information aggregation module 1440 includes:

[0226] The second aggregation module is used to aggregate the content feature information corresponding to the at least two historical published information items to obtain the content aggregation information of the at least two historical published information items.

[0227] The content aggregation information is used to characterize the degree to which the at least two historical published information items contain the same speech recognition information.

[0228] Furthermore, the content feature information aggregation module 1440 includes:

[0229] The averaging processing module is used to average the content feature information corresponding to each historical published information to obtain the central content feature information of each content feature information.

[0230] The similarity calculation module is used to calculate the similarity between the content feature information corresponding to each historical published information item and the central content feature information;

[0231] The third generation module is used to obtain the content aggregation information of the multiple historical releases based on the similarity between the content feature information corresponding to each historical release and the central content feature information.

[0232] Furthermore, the device also includes:

[0233] The first acquisition module is used to acquire the verticality level of multiple annotation objects;

[0234] The first prediction module is used to predict the verticality of the multiple labeled objects by performing verticality prediction based on the domain distribution information and content aggregation information of the multiple labeled objects through a preset verticality fusion model; the preset verticality fusion model includes a first weight corresponding to the domain distribution information and a second weight corresponding to the content aggregation information.

[0235] The weight update module is used to update the first weight and the second weight based on the verticality level of the multiple labeled objects and the predicted verticality of the multiple labeled objects, so as to obtain the target verticality fusion model.

[0236] Furthermore, the weight update module includes:

[0237] The third determining module is used to determine the number of labeled objects at each verticality level based on the verticality level of the multiple labeled objects;

[0238] The sorting module is used to sort the predicted verticality of the multiple labeled objects to obtain the verticality sorting result;

[0239] The fourth determination module is used to determine the predicted object corresponding to each verticality level from the verticality ranking results based on the number of labeled objects at each verticality level.

[0240] The precision and recall determination module is used to determine the precision and recall of the preset verticality fusion model based on the labeled object corresponding to each verticality level and the predicted object corresponding to each verticality level.

[0241] The target verticality fusion model determination module is used to update the first weight and the second weight based on the precision and recall rate to obtain the target verticality fusion model.

[0242] Furthermore, the device also includes:

[0243] The second prediction module is used to predict the verticality of the multiple labeled objects based on the target verticality fusion model, and obtain the target verticality of the multiple labeled objects.

[0244] The fifth determination module is used to determine the verticality level classification threshold based on the target verticality of the labeled object at each verticality level.

[0245] Furthermore, the device also includes:

[0246] The sixth determining module is used to determine the verticality level of the target object based on the verticality of the target object and the verticality level classification threshold;

[0247] The management module is used to manage the published information of the target object based on the verticality level of the target object.

[0248] Furthermore, the area determination module 1410 includes:

[0249] The second acquisition module is used to acquire a set of historical published information of the target object within a preset time period adjacent to the current time.

[0250] The selection module is used to randomly select a preset number of historical published information from the set of historical published information.

[0251] The apparatus provided in the above embodiments can execute the methods provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the above embodiments can be found in the methods provided in any embodiment of this application.

[0252] This embodiment also provides a computer-readable storage medium storing at least one instruction or at least one program, which is loaded by a processor and executed as any of the methods described above in this embodiment.

[0253] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the methods described above.

[0254] This embodiment also provides a device, the structural diagram of which can be found in the following figure. Figure 15 The device 1500 can vary significantly in configuration or performance, and may include one or more central processing units (CPUs) 1522 (e.g., one or more processors) and memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing applications 1542 or data 1544. The memory 1532 and storage media 1530 may be temporary or persistent storage. Programs stored in the storage media 1530 may include one or more modules (not shown), each module including a series of instruction operations on the device. Furthermore, the CPU 1522 may be configured to communicate with the storage media 1530 and execute the series of instruction operations in the storage media 1530 on the device 1500. The device 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM Etc. Any of the methods described above in this embodiment can be based on... Figure 15 The equipment shown is used for implementation.

[0255] This specification provides the operational steps of the methods described in the embodiments or flowcharts, but more or fewer operational steps may be included based on conventional or non-inventive labor. The steps and order listed in the embodiments are merely one possible execution order among many steps and do not represent the only execution order. In actual system or interrupt product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0256] The structure shown in this embodiment is only a partial structure related to the solution of this application and does not constitute a limitation on the device to which the solution of this application is applied. Specific devices may include more or fewer components than shown, or combinations of certain components, or arrangements of different components. It should be understood that the methods, apparatuses, etc., disclosed in this embodiment can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or unit modules through some interfaces.

[0257] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0258] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0259] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for processing published information, characterized in that, include: Obtain multiple historical postings of the target object and determine the domain corresponding to each historical posting; Based on the domain corresponding to each historical release information, domain distribution information corresponding to the target object is determined; the domain distribution information characterizes the degree of dispersion of the domains corresponding to the multiple historical release information items. Feature extraction is performed on each historical published information to obtain the content feature information corresponding to each historical published information; Based on the content feature information corresponding to each historical published information, the multiple historical published information items are aggregated to obtain the content aggregate information of the multiple historical published information items. The content aggregation information characterizes the degree of aggregation among the various historical published information of the target object; The verticality of the target object is obtained by fusing the domain distribution information and the content aggregation information using a target verticality fusion model. The verticality is used to characterize the degree to which the multiple historical posts contain the same posting characteristics; The training method for the target verticality fusion model includes: Get the verticality level of multiple annotation objects; By using a preset verticality fusion model, verticality prediction is performed based on the domain distribution information and content aggregation information of the multiple labeled objects to obtain the predicted verticality of the multiple labeled objects; the preset verticality fusion model includes a first weight corresponding to the domain distribution information and a second weight corresponding to the content aggregation information; Based on the verticality levels of the multiple labeled objects, the number of labeled objects at each verticality level is determined; the predicted verticality of the multiple labeled objects is sorted to obtain a verticality sorting result; based on the number of labeled objects at each verticality level, the predicted objects corresponding to each verticality level are determined from the verticality sorting result. Based on the labeled objects corresponding to each verticality level and the predicted objects corresponding to each verticality level, the precision and recall of the preset verticality fusion model are determined; based on the precision and recall, the first weight and the second weight are updated to obtain the target verticality fusion model.

2. The method according to claim 1, characterized in that, The step of determining the domain distribution information corresponding to the target object based on the domain corresponding to each historical published information includes: Based on the category identifier corresponding to each historical release information, the number of historical release information corresponding to each category identifier is determined; wherein, the domain corresponding to each historical release information is determined by the category identifier corresponding to each historical release information; Based on the number of historical published information corresponding to each category identifier, the domain distribution information corresponding to the target object is determined.

3. The method according to claim 2, characterized in that, The determination of the domain distribution information corresponding to the target object based on the number of historical published information corresponding to each category identifier includes: Based on the number of historical published information corresponding to each category identifier, calculate the category information entropy corresponding to the target object; Based on the category information entropy, the domain distribution information corresponding to the target object is determined.

4. The method according to claim 3, characterized in that, At least two of the aforementioned historical releases include audio information; The step of extracting features from each historical posting to obtain content feature information corresponding to each historical posting includes: Acoustic features are extracted from the audio information of the at least two historical releases to obtain acoustic feature information corresponding to the at least two historical releases. Based on the acoustic feature information, content feature information corresponding to the at least two historical publication information items is obtained; The step of aggregating the multiple historical publications based on the content feature information corresponding to each historical publication to obtain the content aggregation information of the multiple historical publications includes: The content feature information corresponding to the at least two historical published information items is aggregated to obtain the content aggregate information of the at least two historical published information items. The content aggregation information is used to characterize the degree to which the at least two historical published information items contain the same acoustic feature information.

5. The method according to claim 3, characterized in that, At least two of the aforementioned historical releases include audio information; The step of extracting features from each historical posting to obtain content feature information corresponding to each historical posting includes: Perform speech recognition on the audio information of the at least two historical published information items to obtain speech recognition information corresponding to the at least two historical published information items; Based on the speech recognition information, content feature information corresponding to the at least two historical published information items is obtained; The step of aggregating the multiple historical publications based on the content feature information corresponding to each historical publication to obtain the content aggregation information of the multiple historical publications includes: The content feature information corresponding to the at least two historical published information items is aggregated to obtain the content aggregate information of the at least two historical published information items. The content aggregation information is used to characterize the degree to which the at least two historical published information items contain the same speech recognition information.

6. The method according to claim 1, characterized in that, The step of obtaining content aggregation information for the multiple historical published information items based on the content feature information corresponding to each historical published information item includes: The content feature information corresponding to each historical published information is averaged to obtain the central content feature information of each content feature information; Calculate the similarity between the content feature information corresponding to each historical published information item and the central content feature information; Based on the similarity between the content feature information corresponding to each historical published information and the central content feature information, the content aggregation information of the multiple historical published information is obtained.

7. The method according to claim 1, characterized in that, The method further includes: Based on the target verticality fusion model, the verticality of the multiple labeled objects is predicted to obtain the target verticality of the multiple labeled objects. Based on the target verticality of the labeled object at each verticality level, determine the threshold for verticality level division.

8. The method according to claim 7, characterized in that, The method further includes: The verticality level of the target object is determined based on the verticality of the target object and the verticality level classification threshold. Based on the verticality level of the target object, the published information of the target object is controlled.

9. The method according to claim 1, characterized in that, The acquisition of multiple historical publication information of the target object includes: Obtain the set of historical published information of the target object within a preset time period adjacent to the current time; A preset number of historical published information are randomly selected from the set of historical published information.

10. An information publishing and processing device, characterized in that, include: The domain determination module is used to obtain multiple historical publication information of the target object and determine the domain corresponding to each piece of historical publication information. The domain distribution information determination module is used to determine the domain distribution information corresponding to the target object based on the domain corresponding to each historical release information; the domain distribution information characterizes the degree of dispersion of the domains corresponding to the multiple historical release information items; The content feature information generation module is used to extract features from each historical published information to obtain the content feature information corresponding to each historical published information; The content feature information aggregation module is used to aggregate the multiple historical published information items according to the content feature information corresponding to each historical published information item, so as to obtain the content aggregation information of the multiple historical published information items; The content aggregation information characterizes the degree of aggregation among the various historical published information of the target object; The verticality determination module is used to fuse the domain distribution information and the content aggregation information through a target verticality fusion model to obtain the verticality of the target object. The verticality is used to characterize the degree to which the multiple historical posts contain the same posting characteristics; The training method for the target verticality fusion model includes: Get the verticality level of multiple annotation objects; By using a preset verticality fusion model, verticality prediction is performed based on the domain distribution information and content aggregation information of the multiple labeled objects to obtain the predicted verticality of the multiple labeled objects; the preset verticality fusion model includes a first weight corresponding to the domain distribution information and a second weight corresponding to the content aggregation information; Based on the verticality levels of the multiple labeled objects, the number of labeled objects at each verticality level is determined; the predicted verticality of the multiple labeled objects is sorted to obtain a verticality sorting result; based on the number of labeled objects at each verticality level, the predicted objects corresponding to each verticality level are determined from the verticality sorting result. Based on the labeled objects corresponding to each verticality level and the predicted objects corresponding to each verticality level, the precision and recall of the preset verticality fusion model are determined; based on the precision and recall, the first weight and the second weight are updated to obtain the target verticality fusion model.

11. The apparatus according to claim 10, characterized in that, The domain distribution information determination module includes: The quantity determination module is used to determine the quantity of historical published information corresponding to each category identifier based on the category identifier corresponding to each historical published information; wherein, the domain corresponding to each historical published information is determined by the category identifier corresponding to each historical published information; The first determining module is used to determine the domain distribution information corresponding to the target object based on the number of historical published information corresponding to each category identifier.

12. The apparatus according to claim 11, characterized in that, The first determining module includes: The category information entropy calculation module is used to calculate the category information entropy corresponding to the target object based on the number of historical published information corresponding to each category identifier; The second determining module is used to determine the domain distribution information corresponding to the target object based on the category information entropy.

13. The apparatus of claim 12, wherein at least two of the plurality of historical published information items include audio information; The content feature information generation module includes: An acoustic feature extraction module is used to extract acoustic features from the audio information of the at least two historical release information items to obtain acoustic feature information corresponding to the at least two historical release information items; The first generation module is used to obtain content feature information corresponding to the at least two historical publication information based on the acoustic feature information; The content feature information aggregation module includes: The first aggregation module is used to aggregate the content feature information corresponding to the at least two historical published information items to obtain the content aggregation information of the at least two historical published information items. The content aggregation information is used to characterize the degree to which the at least two historical published information items contain the same acoustic feature information.

14. The apparatus according to claim 12, characterized in that, At least two of the aforementioned historical releases include audio information; The content feature information generation module includes: The speech recognition module is used to perform speech recognition on the audio information of the at least two historical release information items to obtain speech recognition information corresponding to the at least two historical release information items; The second generation module is used to obtain content feature information corresponding to the at least two historical publication information based on the speech recognition information; The content feature information aggregation module includes: The second aggregation module is used to aggregate the content feature information corresponding to the at least two historical published information items to obtain the content aggregation information of the at least two historical published information items. The content aggregation information is used to characterize the degree to which the at least two historical published information items contain the same speech recognition information.

15. The apparatus according to claim 10, characterized in that, The content feature information aggregation module includes: The averaging processing module is used to average the content feature information corresponding to each historical published information to obtain the central content feature information of each content feature information. The similarity calculation module is used to calculate the similarity between the content feature information corresponding to each historical published information item and the central content feature information; The third generation module is used to obtain the content aggregation information of the multiple historical releases based on the similarity between the content feature information corresponding to each historical release and the central content feature information.

16. The apparatus according to claim 10, characterized in that, The device further includes: The second prediction module is used to predict the verticality of the multiple labeled objects based on the target verticality fusion model, and obtain the target verticality of the multiple labeled objects. The fifth determination module is used to determine the verticality level classification threshold based on the target verticality of the labeled object at each verticality level.

17. The apparatus according to claim 16, characterized in that, The device further includes: The sixth determining module is used to determine the verticality level of the target object based on the verticality of the target object and the verticality level classification threshold; The management module is used to manage the published information of the target object based on the verticality level of the target object.

18. The apparatus according to claim 10, characterized in that, The domain determination module includes: The second acquisition module is used to acquire a set of historical published information of the target object within a preset time period adjacent to the current time. The selection module is used to randomly select a preset number of historical published information from the set of historical published information.

19. An information publishing and processing device, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the information publishing processing method as described in any one of claims 1 to 9.

20. A computer storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor according to any one of claims 1 to 9.

21. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the information publishing processing method as described in any one of claims 1 to 9.