Data processing methods, apparatus, software products, computer equipment and media

By calculating search frequency fluctuation parameters and outbreak coefficients, and clustering search content, the problem of insufficient accuracy of popular events in existing technologies is solved, and accurate and timely discovery of search outbreak events is achieved.

CN116796077BActive Publication Date: 2026-05-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-03-17
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, determining whether an event is popular solely based on the number of searches ignores the explosive nature of the event, resulting in inaccurate information on popular events.

Method used

By acquiring the set of search content within the target time segment, calculating the frequency fluctuation parameter and search burst coefficient of search frequency, clustering similar content, and mining search burst events associated with the search content.

Benefits of technology

It improves the accuracy of searching for breaking events, ensuring that the events discovered are suddenly trending and popular, thus enhancing the timeliness and accuracy of the events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116796077B_ABST
    Figure CN116796077B_ABST
Patent Text Reader

Abstract

This application discloses a data processing method, apparatus, program product, computer equipment, and medium. The method includes: obtaining a search content set corresponding to a target time segment; the search content set contains search content within multiple time segments of the target time segment, and the search content in the search content set has content similarity; obtaining frequency fluctuation parameters of the search frequency of the search content in the search content set within the multiple time segments of the target time segment; determining the search burst coefficient of the search content in the search content set within the multiple time segments of the target time segment based on the frequency fluctuation parameters; if the search burst coefficient meets the burst coefficient standard, then obtaining search burst events associated with the search content set within the target time segment. Using this application can improve the accuracy of the obtained search burst events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more particularly to a data processing method, apparatus, program product, computer equipment, and medium. Background Technology

[0002] When pushing trending events, recent breaking news can often be identified by analyzing user searches on the platform. For example, events with a high number of searches can be identified as breaking news and then pushed to the platform. However, judging whether an event is a breaking news event solely based on the number of searches ignores the explosive nature of the event, making the identified breaking news events often inaccurate. Summary of the Invention

[0003] This application provides a data processing method, apparatus, program product, computer equipment, and medium that can improve the accuracy of obtained search outbreak events.

[0004] This application provides a data processing method, which includes:

[0005] Obtain the search content set corresponding to the target time segment; the search content set contains search content within the target time segment across multiple time periods, and the search content within the search content set exhibits content similarity.

[0006] Obtain frequency fluctuation parameters of the search frequency of search content in the search content set within target time segments across multiple time periods;

[0007] The search burst coefficient of the search content in the search content set within the target time slices of multiple time periods is determined based on the frequency fluctuation parameter.

[0008] If the search burst coefficient meets the burst coefficient standard, then the search burst events associated with the search content set within the target time slice are obtained.

[0009] This application provides a data processing apparatus, which includes:

[0010] The set acquisition module is used to acquire the search content set corresponding to the target time segment; the search content set contains search content within the target time segment of multiple time periods, and the search content in the search content set has content similarity;

[0011] The parameter acquisition module is used to obtain the frequency fluctuation parameters of the search frequency of the search content in the search content set within the target time segment of multiple time periods;

[0012] The coefficient determination module is used to determine the search burst coefficient of the search content in the search content set within the target time slices of multiple time periods based on the frequency fluctuation parameter.

[0013] The event acquisition module is used to acquire search burst events associated with the search content set within the target time slice if the search burst coefficient meets the burst coefficient standard.

[0014] Optionally, the collection acquisition module may obtain the search content set corresponding to the target time slice in the following ways:

[0015] Retrieve N search results to be processed; N search results contain search results from multiple time periods, where N is a positive integer;

[0016] Cluster the N search results for similar content to obtain at least one category of search results;

[0017] Based on the search content within the target time slice of each search content of at least one type of search content, create a search content set corresponding to each type of search content of at least one type of search content.

[0018] Optionally, the set acquisition module can perform clustering of similar content on N search results to obtain at least one category of search results, including:

[0019] Perform word segmentation on each of the N search terms to obtain the entity words contained in each of the N search terms;

[0020] Normalize and rewrite the entity segments contained in each of the N search terms to obtain the rewritten search terms corresponding to each of the N search terms.

[0021] For each rewritten search result, perform clustering of similar content to obtain at least one category of search results.

[0022] Optionally, the collection acquisition module performs clustering of similar content on each rewritten search result to obtain at least one category of search results, including:

[0023] Generate content embedding features for each rewritten search result;

[0024] Based on the content embedding features of each rewritten search content, clustering of similar content is performed on each rewritten search content to obtain at least one class of search content; any one of the at least one class of search content contains at least one rewritten search content.

[0025] Optionally, the number of rewritten search terms corresponding to N search terms is N, and any one of the N rewritten search terms is represented as the i-th rewritten search term, where i is a positive integer and i is less than or equal to N;

[0026] The collection acquisition module performs clustering of similar content on each rewritten search content based on the content embedding features of each rewritten search content, obtaining at least one class of search content in the following ways:

[0027] Get the M classes of search content clustered before clustering the i-th rewritten search content; each of the M classes of search content contains at least one rewritten search content, and M is a positive integer;

[0028] The content embedding features of the rewritten search content contained in each of the M types of search content are averaged to obtain the average embedding features of the rewritten search content contained in each of the M types of search content.

[0029] Based on the content embedding features of the i-th rewritten search content and the average embedding features corresponding to each of the M types of search content, the i-th rewritten search content is clustered to obtain at least one type of search content.

[0030] Optionally, the set acquisition module performs clustering on the i-th rewritten search content based on the content embedding features of the i-th rewritten search content and the average embedding features corresponding to each of the M types of search content, to obtain at least one type of search content. This includes:

[0031] Obtain the feature distance between the content embedding feature of the i-th rewritten search content and the average embedding feature of each search content in the M-class search content. This feature distance is used to characterize the content similarity between the i-th rewritten search content and the rewritten search content contained in each search content in the M-class search content.

[0032] If there is a target class search content in the M-class search content with a corresponding feature distance less than or equal to the distance threshold, then the i-th rewritten search content is clustered into the target class search content to obtain the clustered target class search content.

[0033] If the target class search content does not exist in the search content of class M, then create a unique class search content based on the i-th rewritten search content;

[0034] Based on the search content of category M, excluding the target category search content after clustering, the target category search content after clustering, or the unique category search content, determine at least one type of search content.

[0035] Optionally, the parameter acquisition module may obtain the frequency fluctuation parameter of the search frequency of search content within multiple target time segments in the search content set in the search content set, including:

[0036] Obtain the search frequency of each search item in the search content set within the target time segment of each time period;

[0037] Based on the search frequency corresponding to each time period, calculate the frequency standard deviation of the search frequency of the search content in the search content set within the target time segment of multiple time periods.

[0038] The frequency standard deviation is defined as the frequency fluctuation parameter.

[0039] Optionally, any one of multiple time periods can be represented as the target time period; the parameter acquisition module obtains the search frequency of the search content in the search content set within the target time segment of each time period in the following ways:

[0040] The search content that falls within the target time segment of the target time period is selected as the search content to be calculated.

[0041] Obtain the search type of each search item to be calculated, and determine the search frequency weight corresponding to each search item to be calculated based on the search type of each search item to be calculated.

[0042] Based on the search frequency weight corresponding to each search content to be calculated, the search frequency of the search content in the search content set within the target time segment of the target time period is determined.

[0043] Optionally, the parameter acquisition module calculates the frequency standard deviation of the search frequency of the search content in the search content set within the target time slices of multiple time periods based on the search frequency corresponding to each time period. This includes methods such as:

[0044] The frequency decay weight for each time period is determined based on the chronological order between each time period.

[0045] The frequency standard deviation is determined based on the frequency decay weight and search frequency corresponding to each time period.

[0046] Optionally, the coefficient determination module determines the search burst coefficient of the search content in the search content set within multiple time segments based on the frequency fluctuation parameter, including:

[0047] Based on the search frequency corresponding to each time period, determine the average search frequency of the search content in the search content set within the target time segments of multiple time periods.

[0048] The search burst coefficient is determined based on the average search frequency and frequency fluctuation parameters.

[0049] Optionally, if the search burst coefficient meets the burst coefficient standard, the event acquisition module obtains search burst events associated with the search content set within the target time slice in the following ways:

[0050] If the search burst coefficient is greater than or equal to the burst coefficient threshold, then the search burst coefficient is determined to meet the burst coefficient standard.

[0051] Retrieve at least one search result associated with the search content set;

[0052] A search outbreak event is determined based on at least one search result.

[0053] Optionally, the event retrieval module may retrieve at least one search result associated with the search content set in the following ways:

[0054] Select search content from the set of search results as a reference;

[0055] At least one search result is obtained based on the search content used as a reference.

[0056] Optionally, the event acquisition module determines the search outbreak event based on at least one search result in the following ways:

[0057] Retrieve the topic text contained in each search result;

[0058] Each topic text is segmented into clauses to obtain multiple text clauses of the topic text contained in at least one search result;

[0059] Obtain the event evaluation index for each text clause, and determine the search burst event from multiple text clauses based on the event evaluation index for each text clause.

[0060] This application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the method of this application.

[0061] This application provides a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method described above.

[0062] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative embodiments described above.

[0063] This application obtains a search content set corresponding to a target time segment; the search content set contains search content within multiple time segments of the target time segment, and the search content in the search content set has content similarity; it obtains the frequency fluctuation parameters of the search frequency of the search content in the search content set within the target time segments of multiple time segments; it determines the search burst coefficient of the search content in the search content set within the target time segments of multiple time segments based on the frequency fluctuation parameters; if the search burst coefficient meets the burst coefficient standard, it obtains the search burst events associated with the search content set within the target time segment. Therefore, the method proposed in this application can synchronously process search content with content similarity in the search content set, and obtain the search burst coefficient of the search content in the target time segment by using the frequency fluctuation coefficient of the search frequency of the search content in the same time segment (such as the target time segment) of each time segment. This takes into account the search burstability (such as volatility) of the search content, and thus obtains the search burst events through the search burstability of the search content, ensuring that the obtained search burst events also have search burstability, thus making the obtained search burst events more accurate. Attached Figure Description

[0064] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0065] Figure 1 This is a schematic diagram of a network architecture provided in this application;

[0066] Figure 2 This is a schematic diagram of an event mining scenario provided in this application;

[0067] Figure 3 This is a flowchart illustrating a data processing method provided in this application;

[0068] Figure 4 This is a schematic diagram illustrating a scenario of target time segmentation provided in this application;

[0069] Figure 5 This is a schematic diagram of an event mining scenario provided in this application;

[0070] Figure 6 This is a schematic diagram of a frequency variation curve provided in this application;

[0071] Figures 7a-7dThis is a schematic diagram of an interface provided in this application for pushing out search outbreak events;

[0072] Figure 8 This is a flowchart illustrating a data processing method provided in this application;

[0073] Figure 9 This is a schematic diagram illustrating a scenario for obtaining rewrite weights provided in this application;

[0074] Figure 10 This is a schematic diagram of an event mining scenario provided in this application;

[0075] Figure 11 This is a schematic diagram of the structure of a data processing device provided in this application;

[0076] Figure 12 This is a schematic diagram of the structure of a computer device provided in this application. Detailed Implementation

[0077] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0078] This application relates to technologies related to artificial intelligence (AI). AI is the theory, methods, technology, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0079] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0080] This application primarily concerns machine learning within artificial intelligence. Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0081] The machine learning involved in this application may include the application of models such as event segmentation models, intent recognition models, and pornography recognition models, as detailed below. Figure 8 The description in the corresponding embodiments.

[0082] This application also relates to blockchain-related technologies. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer. A blockchain consists of a series of blocks sequentially generated in chronological order. Once a new block is added to the blockchain, it cannot be removed. Each block records the data submitted by nodes in the blockchain system. In this application, the mined search burst events can be stored on the blockchain to ensure the immutability of these events.

[0083] This application relates to cloud technology. Cloud technology refers to a managed technology that unifies hardware, software, network, and other resources within a wide area network (WAN) or local area network (LAN) to enable data computation, storage, processing, and sharing.

[0084] Cloud technology is a collective term for network technology, information technology, integration technology, management platform technology, and application technology applied to the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will all require robust system support, which can only be achieved through cloud computing.

[0085] The cloud technology involved in this application mainly refers to the ability of the backend to interact with the frontend through the "cloud," such as the backend being able to push discovered search outbreak events to the frontend for display through the "cloud."

[0086] First, it should be noted that this application may display a prompt interface or pop-up window before and during the collection of user-related data (such as user search content and search time). This prompt interface or pop-up window is used to inform the user that their relevant data is being collected. This application will only begin the steps of collecting user-related data after receiving confirmation from the user regarding the prompt interface or pop-up window; otherwise (i.e., without receiving confirmation from the user), the steps of collecting user-related data will end, meaning no user-related data will be collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of relevant user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0087] Please see Figure 1 , Figure 1 This is a schematic diagram of a network architecture provided in this application. For example... Figure 1 As shown, the network architecture may include server 200 and a cluster of terminal devices. The cluster of terminal devices may include one or more terminal devices; the number of terminal devices is not limited here. Figure 1 As shown, the multiple terminal devices may specifically include terminal device 100a, terminal device 101a, terminal device 102a, ..., terminal device 103a; as Figure 1 As shown, terminal devices 100a, 101a, 102a, ..., 103a can all connect to server 200 via the network, so that each terminal device can interact with server 200 via the network.

[0088] like Figure 1The server 200 shown can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminal devices can be smart terminals such as smartphones, tablets, laptops, desktop computers, and smart TVs.

[0089] Please see also Figure 2 , Figure 2 This is a schematic diagram of an event mining scenario provided in this application. For example... Figure 2 As shown, each terminal device (terminal device 100a, terminal device 101a, terminal device 102a, ..., terminal device 103a) can have a search platform. This search platform can be any platform that can perform information searches (such as a web platform, application platform, etc.). Users holding each terminal device can search for relevant content on the search platform. The content that users search for on the search platform can be called search content. For example, a user's search content can be "Where did an earthquake happen today?" which they enter themselves on the platform.

[0090] Server 200 can be the backend server of the search platform. Therefore, each terminal device can provide the search content obtained by the user on the search platform to Server 200. Server 200 can then mine the popular events of user bursts (such as sudden bursts) based on the obtained search content. The mined bursts of popular events can be called search burst events. Please refer to the following description.

[0091] Server 200 can mine recent search outbreaks by acquiring recent (e.g., the last 3 days) search content. Server 200 can process each recently acquired search content (e.g., cleaning and filtering, and normalizing and rewriting the entity words contained in the search content) to obtain several search contents to be clustered. These search contents to be clustered can be as follows: Figure 8 For the N rewritten search terms in the corresponding embodiment, the specific process of obtaining several search terms to be clustered can be found below. Figure 8 The relevant descriptions in the corresponding embodiments.

[0092] Furthermore, server 200 can cluster several search items to be clustered, grouping search items with the same or similar search intent into the same category, resulting in several categories of search items. These several categories of search items can also be as follows: Figure 8The corresponding embodiment refers to at least one type of search content, and the several types of search content here may include the first type of search content to the fourth type of search content.

[0093] If it is necessary to mine search outbreaks within a target time slice in the recent period (e.g., multiple time periods, where each period can be a recent day), server 200 can extract the search content within the target time slice (e.g., any time interval of a day) from each type of search content to create a search content set corresponding to each type of search content. Here, search content set 1 is created using the search content within the target time slice of the multiple time periods in the first type of search content (the search time is within that target time slice); search content set 2 is created using the search content within the target time slice of the multiple time periods in the second type of search content; search content set 3 is created using the search content within the target time slice of the multiple time periods in the third type of search content; and search content set 4 is created using the search content within the target time slice of the multiple time periods in the fourth type of search content.

[0094] Therefore, server 200 can detect search outbreak events associated with each search content set within the target time segment by analyzing the fluctuations (e.g., the number of searches) of search content in each search content set within the target time segment of the above-mentioned multiple time periods. Then, server 200 can push the detected search outbreak events to the above-mentioned terminal devices for display.

[0095] Using the method provided in this application, search content with the same or similar intent can first be clustered, and then search outbreak events can be mined based on the various types of search content in the clusters. This makes the mined search outbreak events more accurate. In addition, this application takes into account the fluctuation of the search frequency of search content in the search content set within the target time segment of each time period to mine search outbreak events. This ensures that the mined search outbreak events are explosive and suddenly popular events, so as to accurately and timely mine recently popular events.

[0096] Please see Figure 3 , Figure 3 This is a flowchart illustrating a data processing method provided in this application. The executing entity in the embodiments of this application can be a single computer device or a cluster of computer devices, hereinafter collectively referred to as a computer device. This computer device can be a server or a terminal device. Therefore, the executing entity in the embodiments of this application can be a server, a terminal device, or a combination of both. Figure 3 As shown, the method may include:

[0097] Step S101: Obtain the search content set corresponding to the target time segment; the search content set contains search content within the target time segment of multiple time periods, and the search content in the search content set has content similarity.

[0098] Optionally, the computer device can obtain a set of search content corresponding to the target time slice. The set of search content can include search content within the target time slice of multiple time periods. The search content can be the content entered by the user on any search platform, or it can be the content entered by the user on a specific search platform (such as the search platform mainly used for analysis). The search content can be in text form.

[0099] Among the multiple time periods, one of the time periods can be a single calendar day (e.g., one day), or it can be a time period consisting of several consecutive calendar days. Optionally, the specific time period can be set according to the actual time period requirements of the analysis. The following explanation uses a single calendar day as an example.

[0100] The aforementioned multiple time periods can be several calendar days (i.e., several days) prior to the event. For example, the number of time periods can be 5, meaning there are a total of 5 time periods, and these 5 time periods can be the 5 days prior to the event. Since this application is mainly used for mining explosive events, it can mine explosive events (such as events with a sudden surge in search popularity) through real-time search content. Therefore, the multiple time periods in this application can be multiple days prior to the event (including the current day). Optionally, the specific time periods referred to by the multiple time periods can be determined according to the actual application scenario. For example, if the goal is to mine explosive events within the last 10 days, then the multiple time periods can refer to the 10 days prior to the event (including the current day).

[0101] For example, if the above multiple time periods are 5 time periods, and the date for mining the current explosive event is January 6, then these multiple time periods can be from January 2 to January 6, a total of 5 days (i.e. 5 calendar days).

[0102] The aforementioned target time slice can refer to a specific time period within any day. This application can mine explosive events within the target time slice. For example, the duration of the target time slice can be 1 hour (or other durations, such as half an hour). Therefore, the target time slice can be from 0:00 to 1:00, 1:00 to 2:00, ..., or from 23:00 to 24:00 within a day; that is, the target time slice can be any hour within a day. Furthermore, if explosive events are mined using a sliding window approach, with the sliding unit being 1 minute and the target time slice duration being 1 hour, then the target time slice can refer to from 0:00 to 1:00, 0:01 to 1:01, 0:02 to 1:02, ..., or from 23:00 to 24:00 within a day. Since the principle of mining relevant explosive events is the same in any time interval (such as any time period within a day), the following explanation will take mining explosive events within a target time segment as an example. The target time segment can be any time period within a day, and the duration of the target time segment can also be arbitrary, depending on the actual application scenario.

[0103] Therefore, the aforementioned search content set can include search content within the same time slice (i.e., target time slice) of the aforementioned multiple time periods. In other words, the search time of the search content in the search content set falls within the target time slice of those multiple time periods. For example, these multiple time periods can include three time periods: January 1st, January 2nd, and January 3rd (one day is one time period). The target time slice can be 10:00-10:30. Therefore, the target time slices for these multiple time periods can include 10:00-10:30 on January 1st, 10:00-10:30 on January 2nd, and 10:00-10:30 on January 3rd.

[0104] Furthermore, the various search items within the aforementioned search content set also exhibit content similarity. This content similarity indicates that the search intent of each search item in the set is the same or similar. For details on how to obtain the search content set, please refer to the following... Figure 8 The relevant descriptions in the corresponding embodiments.

[0105] Please see Figure 4 , Figure 4 This is a schematic diagram illustrating a scenario of target time segmentation provided in this application. For example... Figure 4As shown, if the above multiple time periods include January 1st, January 2nd, and January 3rd, a total of 3 days, and the current time is 9:30 on January 3rd, and the search burst events in the first half hour are mined at the current time, then the target time slice can be 9:00 to 9:30 of the day. Therefore, the target time slices for multiple time periods can include 9:00 to 9:30 on January 1st, 9:00 to 9:30 on January 2nd, and 9:00 to 9:30 on January 3rd.

[0106] If we mine explosive events (such as search outbreaks described below) in real time using a sliding window, with each slide lasting one minute, then the next target time slice after the current time (e.g., 9:30 AM on January 3rd) would be the target time slice that is slid forward one minute, which is 9:01 AM on the same day. Therefore, target time slices for multiple time periods could include 9:01-9:31 AM on January 1st, 9:01-9:31 AM on January 2nd, and 9:01-9:31 AM on January 3rd. As time progresses, the target time slices can be continuously updated. The principle for mining search outbreaks associated with search content sets within each target time slice is the same, as described below. This application uses the process of mining search outbreaks associated with search content sets within a single target time slice as an example.

[0107] Step S102: Obtain the frequency fluctuation parameters of the search frequency of the search content in the search content set within the target time segments of multiple time periods.

[0108] Optionally, the computer device can obtain frequency fluctuation parameters of the search frequency of the search content in the search content set within the target time segments of the aforementioned multiple time periods. These frequency fluctuation parameters ensure the fluctuation range of the search frequency of the search content in the search content set between the target time segments of each time period. A larger frequency fluctuation parameter results in a larger fluctuation range, and vice versa. The calculation process for these frequency fluctuation parameters can be described as follows.

[0109] First, the computer device can obtain the search frequency of the search content in the search content set within the target time segment of each of the above multiple time periods. This process can be as follows: any one of the multiple time periods can be represented as the target time period. Here, we take obtaining the search frequency of the search content in the search content set within the target time segment of the target time period as an example. It can be understood that the principle of obtaining the search frequency of the search content in the search content set within the target time segment of each time period is the same.

[0110] Computer devices can focus search content within a target time segment of a target time period as the search content to be calculated. Therefore, the computer device can obtain the search type of each search content to be calculated. Optionally, the search type can include two types: active search and passive search. Active search content refers to content that the user manually enters into the search platform (e.g., a user-entered search query like "What magnitude earthquake occurred in xxx city?"). Passive search content can refer to content provided by the search platform that allows users to click and search for relevant information (e.g., keywords displayed on the search platform or terms in trending searches).

[0111] Because active search results are based on the user's subjective search, while passive search results are based on recommendations from the search platform, active search results can have a higher weight than passive search results. This weight can be called search frequency weight. For example, the search frequency weight for active search results could be 1.5, while the search frequency weight for passive search results could be 1.

[0112] Therefore, the computer device can obtain the search frequency weight corresponding to each search item to be calculated based on the search type of each search item. Then, the computer device can obtain the search frequency of the search items in the search content set within the target time segment of the target time period based on the search frequency weight corresponding to each search item. For example, the search frequency weights corresponding to each search item to be calculated can be added together, and the sum can be used as the search frequency of the search items in the search content set within the target time segment of the target time period.

[0113] For example, each search content to be calculated may include search content 1, search content 2, search content 3, and search content 4. Among them, search content 1 is an active search type search content, so the search frequency weight corresponding to search content 1 can be 1.5. Search content 2 can also be an active search type search content, so the search frequency weight corresponding to search content 2 can also be 1.5. Search content 3 is a passive search type search content, so the search frequency weight corresponding to search content 3 can be 1. Search content 4 can also be a passive search type search content, so the search frequency weight corresponding to search content 4 can also be 1. Therefore, the search frequency of the search content in the search content set within the target time segment of the target time period can be equal to 1.5 + 1.5 + 1 + 1, which equals 5.

[0114] Through the above process, the computer device can calculate the search frequency of the search content in the search content set within the target time segment of each time period, with one time period corresponding to one search frequency.

[0115] Furthermore, computer equipment can calculate the standard deviation of the search frequency of the search content in the search content set within the target time segment of multiple time periods based on the search frequency corresponding to each time period. This standard deviation can be called the frequency standard deviation, and then it can be used as the frequency fluctuation parameter mentioned above.

[0116] Since the standard deviation in the general sense refers to the arithmetic square root of the variance, in calculating the above frequency standard deviation in this application, the timeliness of each search content can be taken into account. The search content that is further away from the present can have a smaller weight, and the search content that is closer to the present can have a larger weight. This weight can be called the time decay weight.

[0117] Therefore, the computer device can obtain the time decay weight corresponding to each time period. The further back in time a time period is from the present, the smaller its time decay weight; the closer the time period is to the present, the larger its time decay weight. The computer device can determine the time decay weight corresponding to each time period based on the chronological order of the time periods. For example, firstly, a decay coefficient (decay) less than 1 can be defined, and the specific value of this decay coefficient (decay) can be set according to the actual application scenario.

[0118] Therefore, the time periods can be sorted chronologically from late to early (i.e., the closest time period is listed first, and the furthest time period is listed last). Thus, the time decay weight corresponding to the first time period can be equal to decay (i.e., decay to the power of 1), and the time decay weight corresponding to the second time period can also be equal to decay. 2 (i.e., decay squared), the time decay weight corresponding to the third time interval can be equal to decay. 3 (i.e., decay cubed), and so on. Assuming there are n time periods in total, the time decay weight corresponding to the nth time period can be equal to decay. n (i.e., decay raised to the power of n).

[0119] Furthermore, computer equipment can also calculate the average search frequency (i.e., average value) of the search content within a target time segment across multiple time periods, based on the search frequencies corresponding to each time period. For example, the aforementioned multiple time periods may include time period 1, time period 2, and time period 3. If the search frequency for time period 1 is 10, the search frequency for time period 2 is 50, and the search frequency for time period 3 is 30, then the average search frequency can be equal to... That is, it equals 30.

[0120] Therefore, the computer device can obtain the above frequency standard deviation based on the average search frequency, the frequency decay weight corresponding to each time period, and the search frequency. This frequency standard deviation can be denoted as weight.sd, and the process can be shown in the following formula (1):

[0121]

[0122] Here, avg represents the average search frequency, x1 represents the search frequency corresponding to the first time period ordered chronologically, x2 represents the search frequency corresponding to the second time period ordered chronologically, and so on. n This can represent the search frequency corresponding to the nth time period when the above items are sorted in chronological order.

[0123] This application takes into account the time decay weights corresponding to each time period based on their proximity, and then uses these time decay weights to calculate the frequency standard deviation of the search content in the target time segment of each time period. The calculated frequency standard deviation is more accurate and more in line with actual application scenarios.

[0124] Step S103: Determine the search burst coefficient of the search content in the target time slice of multiple time periods based on the frequency fluctuation parameter.

[0125] Optionally, the computer device can obtain the search burst coefficient of the search content in the search content set within the target time segment in multiple time periods based on the frequency fluctuation parameters calculated above. The search burst coefficient can be used to measure whether the search content in the search content set has reached a burst state within the target time segment.

[0126] It is understandable that, since the search content in the search content set has similarity, it can be understood that the search content in the search content set belongs to the same type of search content. Therefore, the search burst coefficient can also be used to characterize whether the search content of this type belongs to the burst search content within the target time segment, that is, whether it belongs to the burst search content that suddenly becomes popular.

[0127] The above search burst coefficient can be expressed as weight.qcv, and the calculation process of weight.qcv can be shown in the following formula (2):

[0128]

[0129] Here, avg can represent the average search frequency mentioned above, weight.sd represents the frequency fluctuation parameter mentioned above, and const is a small constant (such as a constant greater than 0 but close to 0) used to prevent calculation errors in extreme cases (such as avg equal to 0).

[0130] Step S104: If the search burst coefficient meets the burst coefficient standard, then obtain the search burst events associated with the search content set within the target time slice.

[0131] Optionally, if the search burst coefficient is greater than or equal to the burst coefficient threshold, it can be considered that the search burst coefficient meets the burst coefficient standard. This indicates that the search content in the search content set has reached a burst state within the target time segment, which means that the search content in the search content set is a burst of suddenly popular search content within the target time segment. At this time, the computer device can discover the events associated with the search content in the search content set and regard the event as the search burst event. The search burst event is a popular event with a sudden increase in search volume within the target time segment over multiple time periods, i.e., a sudden popular event.

[0132] Optionally, the process of obtaining search outbreak events associated with the search content set can be:

[0133] First, the computer device can obtain at least one search result associated with the search content set. This process can be as follows: the computer device can obtain search content associated with the search content set for reference. Optionally, the search content used for reference can be any search content in the search content set, or the search content used for reference can be any search content in the category of search content to which the search content set belongs. The search content included in the category of search content to which the search content set belongs all possess content similarity. The process of obtaining the category of search content to which the search content set belongs can be described below. Figure 8 The relevant description in the corresponding embodiment. Therefore, the computer device can input the search content used for reference into a search platform (or other platforms), and obtain at least one search result by searching the search content used for reference on the search platform. Any search result can be a popular science result, an article result, a discussion result, or a question and answer result, etc.

[0134] Therefore, the computer device can determine the search outbreak event using at least one search result obtained above: Optionally, the computer device can obtain the topic text contained in each search result, which may be the title text contained in the search result. Furthermore, the computer device can segment the obtained topic text into clauses, obtaining clauses (which can be called text clauses) contained in each topic text. A topic text can be segmented into one or more text clauses.

[0135] Next, the computer device can obtain the event evaluation index for each text clause, which indicates the probability that the corresponding text clause describes an event. Therefore, the computer device can determine search outbreak events from all text clauses contained in each topic text based on the event evaluation indices of each text clause. For example, the computer device can identify the text clause with the highest event evaluation index and / or whose highest event evaluation index is greater than an evaluation index threshold as the search outbreak event. That is, the event title of this search outbreak event can be this text clause (e.g., the text clause with the highest event evaluation index, or the text clause with the highest event evaluation index and whose highest event evaluation index is greater than an evaluation index threshold, or the text clause with an event evaluation index greater than an evaluation index threshold). The event described by this text clause is the search outbreak event.

[0136] The event evaluation index of each text clause can be obtained by a pre-trained event segmentation model. This event segmentation model can be a pre-trained model used to evaluate the probability (i.e., the score, i.e. the event evaluation index) that the input text describes an event. Therefore, the event evaluation index of each text clause can be obtained by inputting each of the above text clauses into this event segmentation model.

[0137] Optionally, in addition to obtaining the event title of a search outbreak event from several text clauses through the event evaluation index of each text clause, the event title of a search outbreak event can also be obtained from several text clauses through the clause clustering weight and length weight of each text clause, as described below.

[0138] In typical scenarios of pushing trending events, the longer the text clause (e.g., the more words), the more likely it is to describe an event. Therefore, the longer a text clause is, the greater its length weight can be. The method for determining the length weight can be set according to the actual application scenario.

[0139] Furthermore, the text clauses can be clustered to obtain at least one class of text clauses. Therefore, the clause clustering weight of each text clause can be equal to the number of all text clauses contained in its respective class. The process of clustering the text clauses to obtain at least one class of text clauses is as follows: Figure 8 The process of clustering the N search results to obtain at least one category of search results in the corresponding embodiment is the same.

[0140] Therefore, the event evaluation index, length weight, and clause clustering weight corresponding to each text clause can be summed (or weighted and summed again) to obtain the final event evaluation parameters of each text clause. Then, the text clause with the largest event evaluation parameter can be used as the event title of the mined search outbreak event.

[0141] Please see Figure 5 , Figure 5 This is a schematic diagram of an event mining scenario provided in this application. For example... Figure 5 As shown, ① this application can recall resources, where the recalled resources can refer to recalling at least one search result associated with the search content set mentioned above; ② this application can segment the title of the recalled resources (such as the topic text mentioned above) to obtain several clauses of the title of each recalled resource (i.e., the text clauses mentioned above); ③ further, this application can aggregate (i.e., cluster) the clauses, and the aggregation method is as follows: Figure 8 In the corresponding embodiments, the aggregation method for the N search contents to be processed is the same; ④ This application can select the optimal clause (such as the text clause with the largest event evaluation parameter mentioned above) based on the result of clause clustering; ⑤ This application can finally generate the event title of the search outbreak event mined by the selected optimal clause.

[0142] Since there can be multiple search content sets, and search content in different search content sets belongs to different categories (e.g., different intents), while search content in the same search content set belongs to the same category (e.g., same intent), search outbreak events associated with each search content set can be mined independently through the above process. Therefore, one or more search outbreak events can be mined from multiple search content sets. The computer device can also arrange the mined one or more search outbreak events in reverse order (e.g., the more popular the search outbreak event (e.g., the search outbreak event with the higher the search outbreak coefficient of the search content in the associated search content set) can be arranged first), and then push the sorted search outbreak events to the front end (e.g., user terminals with search platforms) for display, such as displaying them in the hot search list, so that users of the search platform can learn about the recent suddenly popular and explosive events (i.e., search outbreak events) in real time and quickly.

[0143] Please see Figure 6 , Figure 6 This is a schematic diagram of a frequency variation curve provided in this application. For example... Figure 6 As shown, if a certain type of search content (such as search content within a certain search content set) has no associated search outbreak events (i.e., no sudden hot events), then the search frequency of this type of search content at every moment can be represented by curve 1. If a certain type of search content (such as search content within a certain search content set) is associated with a sudden search outbreak event, then the search frequency of this type of search content at every moment can be represented by curve 2 (the search frequency was relatively low before, and has recently increased sharply). After a certain type of search content (such as search content within a certain search content set) is associated with a sudden search outbreak event, the search frequency of this type of search content at every moment can be represented by curve 3 (the search frequency was relatively low before, increased sharply during the outbreak, and then decreased again after the outbreak). Therefore, by considering the fluctuations in the search frequency of search content within the target time segment of each time period in this application (which can be reflected by the frequency fluctuation parameter) to mine search outbreak events, it is possible to promptly and quickly discover recently sudden hot events.

[0144] Optionally, the search outbreaks discovered in this application can be displayed in the following product modules:

[0145] 1. Hot Search Ranking: This product sorts all the discovered events (such as search explosion events) in reverse order of popularity, generating an authoritative list of the hottest events across the entire network. The main product entry points include direct access, start page, and homepage.

[0146] 2. Local Ranking: This product categorizes all discovered events according to geographical tags and generates corresponding local hot rankings for different cities. The entry point is the same as the hot search ranking.

[0147] 3. Timely Highlights: For high-profile events (such as search surge events), a timely highlight will be generated at the top of the search results page, including the header image and selected high-quality articles related to the event.

[0148] 4. Event Context: For events with multiple reversals or many subsequent developments, an event context will be used to connect the events at each node (such as a search outbreak event), making it easier for users to quickly grasp the causes and consequences of the event and the latest developments.

[0149] 5. Featured Events in the Feed: The main feed on the homepage will feature some authoritative and trending events (such as search explosion events), along with their corresponding graphic and text resources.

[0150] 6. Keyword Recommendation: Some high-profile events (such as search outbreaks) will be extracted from the homepage and the direct search box as recommended search terms.

[0151] Please see Figures 7a-7d , Figures 7a-7d This is a schematic diagram of an interface provided in this application for pushing out search outbreak events. For example... Figure 7a As shown, this application can display the discovered search outbreak events in the interface 100c of the overall network hot search list, and the content in box 102c indicates that interface 100c is the overall network hot search list; or, this application can also display the discovered search outbreak events in the interface 101c of the local hot search list, and the content in box 103c indicates that interface 101c is the Shenzhen local hot search list.

[0152] For example Figure 7b As shown, this application can also display the discovered search burst events in the interface 104c of the failed large card. The search burst events displayed in interface 104c can be the events in box 105c. For example... Figure 7c As shown, this application can also display the discovered search outbreak events in the event context interface 106c, which can be one or more events in the event context.

[0153] For example Figure 7d As shown, this application can also display the discovered search outbreak events in the interface where hot searches are attached to the information feed (as shown in box 108c), or it can also display the discovered search outbreak events in the interface area of ​​the recommended keywords (as shown in box 107c).

[0154] Furthermore, if the search burst coefficient is less than the aforementioned burst coefficient threshold, meaning the search burst coefficient does not meet the burst coefficient standard, it indicates that the search content in the search content set is not in a burst state, and the events associated with the search content set are not burst-related hot events. Therefore, there is no need to obtain the events associated with the search content set, and no burst-related hot events associated with the search content set have been discovered.

[0155] This application obtains a search content set corresponding to a target time segment; the search content set contains search content within multiple time segments of the target time segment, and the search content in the search content set has content similarity; it obtains the frequency fluctuation parameters of the search frequency of the search content in the search content set within the target time segments of multiple time segments; it determines the search burst coefficient of the search content in the search content set within the target time segments of multiple time segments based on the frequency fluctuation parameters; if the search burst coefficient meets the burst coefficient standard, it obtains the search burst events associated with the search content set within the target time segment. Therefore, the method proposed in this application can synchronously process search content with content similarity in the search content set, and obtain the search burst coefficient of the search content in the target time segment by using the frequency fluctuation coefficient of the search frequency of the search content in the same time segment (such as the target time segment) of each time segment. This takes into account the search burstability (such as volatility) of the search content, and thus obtains the search burst events through the search burstability of the search content, ensuring that the obtained search burst events also have search burstability, thus making the obtained search burst events more accurate.

[0156] Please see Figure 8 , Figure 8 This is a flowchart illustrating a data processing method provided in this application. The execution entity in the embodiments of this application can be the same as described above. Figure 3 The executing entities in the corresponding embodiments are the same, such as computer devices. Figure 8 As shown, the method includes:

[0157] Step S201: Obtain N search terms to be processed; the N search terms contain search terms to be processed within multiple time periods, where N is a positive integer.

[0158] Optionally, the computer device can obtain several original search results (which can be referred to as several original search results) of all users in the search platform within the multiple time periods. For example, the multiple time periods can be the previous 5 days, and the several original search results can include the search results of all users in the search platform within the 5 days (i.e., the search results within the 5 days; it is possible that the 5th day (i.e. the current day) has not yet ended, so the search results for the 5th day can include the search results that already exist on that day).

[0159] Therefore, computer devices can filter several original search results from multiple time periods to obtain N search results to be processed, as described below.

[0160] Filtering each original search content includes two parts. One part is to perform structural filtering on each original search content, and the other part is to perform semantic filtering on the search content after structural filtering:

[0161] First, the structural filtering of each original search content can include: performing word segmentation on each original search content separately to obtain the word segmentation results of each original search content. The word segmentation result corresponding to an original search content can include multiple segmented words obtained by splitting the original search content. Therefore, the computer device can clean and filter empty search contents (search contents without characters), single-character search contents (because usually single-character search contents do not describe events), special word segmentations (which can be filtered according to ASCII (American Standard Code for Information Interchange) and special character tables, mainly referring to some special symbols), word segmentations of useless word types (such as auxiliary words, adjectives, and adverbs, etc.), word segmentations belonging to stop words (such as conjunctions and pronouns, etc.), and word segmentations with too low weights in all original search contents, so as to obtain the search contents after structural filtering corresponding to each original search content respectively.

[0162] Among them, the weights of each segmented word included in each original search content can be obtained through a pre-trained weight model. The lower the weight of a segmented word, the less important the segmented word indicates (for example, segmented words that appear very frequently are usually less important, such as the segmented word "de"). On the contrary, the higher the weight of a segmented word, the more important the segmented word indicates. Therefore, the weights of each segmented word included in each original search content can be obtained according to this weight model, and the segmented words with too low weights (such as lower than a certain weight threshold) can be filtered.

[0163] After performing structural filtering on each original search content separately, the remaining parts included in each original search content after structural filtering can be combined in the original order to form the search content after structural filtering. One original search content corresponds to one search content after structural filtering. For example, one of several original search contents can be "What festival is today". The word segmentation result of this original search content "What festival is today" can include the segmented words "today", "is", "what", and "festival". The remaining part after filtering this original search content can include the segmented words "today" and "festival". Therefore, the search content after structural filtering corresponding to this original search content is the result of splicing the segmented words "today" and "festival", that is, "today festival".

[0164] Then, semantic filtering is performed on each search content after structural filtering:

[0165] Computer devices can also filter meaningless search content (such as randomly piled-up Chinese characters), non-compliant search content (such as search content with adverse social impact, such as segmented words with pornographic or violent attributes), and search content that describes things that are not timely or have no event intent (such as non-timely and non-event-related segmented words in novels and movies).

[0166] Optionally, the computer device can generate quality scores for the filtered search content using a pre-trained quality model. Search content with lower quality scores is considered less meaningful (e.g., randomly piled-up Chinese characters, describing little of the actual event), while search content with higher quality scores is considered more meaningful (describing a more substantial event). Therefore, search content with excessively low quality scores (e.g., below a certain quality threshold) can be considered meaningless and filtered out.

[0167] Optionally, the computer device can use a pre-trained pornography recognition model to determine the probability that the search content after filtering each structure belongs to pornography. Search content with a high probability (such as exceeding a certain threshold) can be considered as non-compliant search content and filtered out.

[0168] Similarly, computer devices can also use pre-trained brute-force identification models to determine the probability that the search content after filtering each structure belongs to brute-force content. Search content with a high probability (such as exceeding a certain threshold) can be considered as non-compliant search content and filtered out.

[0169] Optionally, the computer device can also identify the event intent of the search content after filtering by a pre-trained intent recognition model, and filter out search content whose event intent is not time-sensitive or has a low attribute of being an event.

[0170] The above process yields the filtered search results for all search content within multiple time periods (such as the original search content mentioned above). These filtered search results are the search results to be processed below. The number of these filtered search results can be N, where N is a positive integer. The specific value of N can be determined based on the actual application scenario.

[0171] Step S202: Perform clustering of similar content on the N search results to obtain at least one category of search results.

[0172] Optionally, the computer device can perform word segmentation on each of the N search contents to be processed to obtain entity words contained in each of the N search contents. The entity words can be entity words (such as pronouns, nouns, etc.) in the search contents. The entity words obtained by segmenting the search contents to be processed can be called entity words.

[0173] Next, the computer device can normalize and rewrite the entity segments contained in the N search contents obtained above to obtain the rewritten search contents corresponding to the N search contents. Any one of the N search contents can correspond to a rewritten search content.

[0174] This is because several search terms among N search terms express the same intent, but may use different words (or have significant differences). Therefore, by rewriting several search terms among N search terms, entity segments with the same meaning but different words can be uniformly rewritten into the same entity segments. Then, based on the text similarity between the rewritten N search terms (which can be reflected by the feature distance between the content embedding features of the search terms below), the rewritten N search terms can be clustered. The clustering results (such as at least one category of search terms obtained by clustering below) will be more accurate and more in line with actual application scenarios.

[0175] Specifically, a knowledge graph can be used to normalize and rewrite the entity segments contained in the N search contents. This knowledge graph can contain the relationships between several entities (entities can refer to entity words, and these entities can contain the entity segments contained in the N search contents). Entities with relationships can be connected sequentially according to the context.

[0176] For example, a knowledge graph may include the entity words "China's capital", "Beijing", and "Yandu". In fact, these three entity words describe the same city and are related to each other. Therefore, in the knowledge graph, there can be a connection between "China's capital" and "Beijing", and there can also be a connection between "Beijing" and "Yandu". "China's capital", "Beijing", and "Yandu" are entity words on a relation chain. Therefore, if the entity words contained in the above N search results include these three entity words, then "China's capital" can be rewritten as "Beijing", and "Yandu" can also be rewritten as "Beijing" to achieve normalized rewriting, such as rewriting them all as "Beijing".

[0177] Therefore, this application can normalize and rewrite the entity segments contained in N search contents using a pre-constructed knowledge graph. The specific rule for rewriting an entity word on a relationship chain to a particular entity word can be pre-defined. For example, this rule could include rewriting an entity word to the thing it describes, such as rewriting "China's capital" to "Beijing" itself.

[0178] The above process yields the entity segmentation and rewriting results for each of the N search terms.

[0179] Furthermore, the computer device can perform clustering processing on the various rewritten search results to obtain at least one category of search results. The search results within the same category possess content similarity, which indicates that the search intent of the corresponding search results is the same or similar.

[0180] Optionally, the process by which the computer device clusters the various rewritten search results can be:

[0181] Computer devices can generate content embedding features for each rewritten search content. For example, each rewritten search content can be input into a trained BERT model (a word vector model), and the BERT model can generate content embedding features for each rewritten search content. These content embedding features can be feature vectors of the rewritten search content (or they can be represented as matrices).

[0182] In this application, when generating the content embedding features of each rewritten search content using the BERT model, it is not necessary to generate them based on the frequency of word occurrences in each rewritten search content. Instead, they can be generated by feature learning of each rewritten search content. The generated content embedding features of the rewritten search content can be 128-dimensional embeddings (embedding vectors).

[0183] Furthermore, when generating the content embedding features of each rewritten search content, this application also considers the rewriting weight of each rewritten search content. This rewriting weight can be obtained through the above-mentioned normalization rewriting. The connections between entity words in the knowledge graph can have connection weights. The rewriting weight of a rewritten entity word can be equal to the sum of the weights of all connections from that entity word to another entity word rewritten by that entity word. If a search content contains multiple rewritten entity words, then the rewriting weight of the search content is equal to the sum of the rewriting weights corresponding to the multiple rewritten entity words.

[0184] Therefore, when inputting the rewritten search terms into the BERT model to generate corresponding content embedding features, the rewriting weights corresponding to each rewritten search term can also be input together. This allows the BERT model to consider the rewriting weights of each rewritten search term when generating content embedding features. This ensures that when two rewritten search terms are similar but have significantly different rewriting weights, the content embedding features generated by the BERT model will also have some differences. The greater the difference (e.g., the larger the difference in rewriting weights) between the two rewritten search terms, the greater the difference in the generated content embedding features (indicating that the two rewritten search terms before rewriting also had some differences). Conversely, the smaller the difference in rewriting weights, the smaller the difference in the generated content embedding features.

[0185] Please see Figure 9 , Figure 9 This is a schematic diagram illustrating a scenario for obtaining rewrite weights provided in this application. For example... Figure 9 As shown, a simplified version of a knowledge graph is illustrated. This knowledge graph includes entity words 1 to 7. Entity words with relationships are connected by lines (i.e., edges). Each line has a weight indicating the association between two entity words (which can be called the line weight).

[0186] Therefore, if entity word 6 is rewritten to entity word 7, the rewriting weight of entity word 6 can be equal to 0.7 + 0.3, which equals 1. Similarly, if entity word 3 is rewritten to entity word 4, the rewriting weight of entity word 3 can be equal to 0.6.

[0187] Furthermore, this application can incorporate a large amount of prior information into the BERT model to learn and generate content embedding features for each rewritten search content. This large amount of prior information can be the title (similar to topic text) information in the search results obtained through each rewritten search content. Since search content often has a strong correlation with its corresponding search results, this application can also assist in generating content embedding features for the rewritten search content by having the BERT model learn the features of the title information in the search results corresponding to the rewritten search content. This makes the generated content embedding features of the rewritten search content similar to the embedding features of the title information in its corresponding search results. Therefore, by introducing prior information from the title information in the search results corresponding to the rewritten search content, the generated content embedding features of the rewritten search content can be made more accurate.

[0188] Subsequently, the computer device can cluster the various rewritten search contents using the content embedding features of each rewritten search content to obtain at least one of the aforementioned search contents, as described below.

[0189] There can be a total of N rewritten search terms. Any one of these N rewritten search terms can be represented as the i-th rewritten search term, where i is a positive integer and i is less than or equal to N.

[0190] Since the method of classifying (i.e. clustering) each of the rewritten search contents (which may not be included in the first rewritten search content taken from the N rewritten search contents) is the same, the process of clustering the i-th rewritten search content will be used as an example for explanation.

[0191] Optionally, clustering can be performed by sequentially extracting (e.g., extracting in chronological order or randomly) one rewritten search term from the N rewritten search terms. When extracting the first rewritten search term from the N rewritten search terms, since no other rewritten search terms have been extracted yet, a new category of search terms can be created directly based on the first extracted rewritten search term. In other words, the first extracted rewritten search term is directly classified into a category, and this category of search terms only contains the first extracted rewritten search term.

[0192] Therefore, it is understood that the aforementioned i can be greater than or equal to 2, meaning that the i-th rewritten search content does not include the first rewritten search content extracted, and the i-th rewritten search content can refer to the i-th rewritten search content extracted. Therefore, it is understood that M classes of search content clustered before clustering the i-th rewritten search content can be obtained, where M is a positive integer. Any class of the M search content contains at least one extracted rewritten search content, and the search content contained in any class of the M search content possesses content similarity. In other words, in this application, only search content that possesses content similarity (such as having the same or similar representational intent) can be grouped into one class.

[0193] The computer device can average the content embedding features of the rewritten search content contained in each of the M types of search content to obtain the average embedding features of the rewritten search content contained in each of the M types of search content. Each type of search content in the M types of search content has an average embedding feature (which can also be called the central feature of each type of search content).

[0194] For example, any category of search content in category M includes rewritten search content 1, rewritten search content 2, and rewritten search content 3. The content embedding feature of search content 1 is (1, 1, 1), the content embedding feature of search content 2 is (2, 2, 2), and the content embedding feature of search content 3 is (3, 3, 3). Therefore, the average embedding feature of this category of search content can be obtained by averaging the feature values ​​at each position. For example, the average embedding feature of this category of search content could be... That is, (2, 2, 2). Here, we are only using a 3-dimensional content embedding feature as an example for illustration. In reality, the content embedding feature can be 128-dimensional or 512-dimensional, etc., depending on the actual application scenario.

[0195] Then, the computer device can perform clustering on the i-th rewritten search content based on the content embedding features of the i-th rewritten search content and the average embedding features of each type of search content in the M types of search content, until all N rewritten search content are extracted and clustered separately (i.e., from i equals 2 to i equals N), thus obtaining at least one of the above-mentioned search content.

[0196] Optionally, the process by which the computer device performs clustering on the i-th rewritten search content based on the content embedding features of the i-th rewritten search content and the average embedding features of each of the M types of search content can be as follows:

[0197] The computer device can obtain the feature distances between the content embedding features of the i-th rewritten search content and the average embedding features corresponding to each search type in the M classes of search content. The i-th rewritten search content has a feature distance with any of the M classes of search content; this feature distance can be the cosine distance (e.g., Euclidean distance) between feature vectors. The feature distance between the i-th rewritten search content and any class of search content represents the content similarity between the i-th rewritten search content and all the search content contained in that class. The smaller the feature distance, the greater the content similarity; conversely, the larger the feature distance, the smaller the content similarity.

[0198] Search content whose corresponding feature distance is less than or equal to a distance threshold can be referred to as target search content. All search content included in this target search content class shares content similarity with the i-th rewritten search content. The distance threshold can be set according to the actual application scenario and is not restricted.

[0199] Therefore, if the target class search content exists in the search content of class M (usually only one exists), the i-th rewritten search content can be clustered into the target class search content to obtain the clustered target class search content, which in turn contains the i-th rewritten search content.

[0200] If the target search content does not exist in the M-class search content, that is, if any search content in the M-class search content does not have content similarity with the i-th rewritten search content, then a new type of search content can be created based on the i-th rewritten search content. This new type of search content can be called a unique type of search content. At this time, the unique type of search content only contains the i-th rewritten search content.

[0201] Then, the computer device can continue to cluster the subsequently retrieved rewritten search content using the M-class search content, excluding the clustered target class search content, the clustered target class search content, or the unique class search content, until i equals N, that is, until all the rewritten search content is clustered, thus obtaining at least one of the above-mentioned search content classes.

[0202] For example, if the target class search content exists in the M-class search content when clustering the i-th rewritten search content, then the target class search content will also exist after this clustering process, and the unique class search content created will not exist. Therefore, when clustering the next retrieved search content, the target class search content after clustering and the target class search content after clustering in the M-class search content can be used as the M-class search content for clustering the next retrieved rewritten search content.

[0203] Therefore, following the above, if the target class search content is not present in the M-class search content when clustering the i-th rewritten search content, then there will be no target class search content after this clustering process. After clustering the i-th rewritten search content, when clustering the next retrieved search content, the M-class search content corresponding to the i-th rewritten search content, as well as the unique class search content created based on the i-th rewritten search content, can all be used as the M-class search content for clustering the next retrieved rewritten search content (at this time, the i mentioned above can be increased by 1 to become the new i). The M at this time (the time when the i mentioned above is increased by 1 to become the new i) is increased by 1 compared to the previous time (that is, the time when clustering the i-th rewritten search content).

[0204] From the above, it can be understood that the number of M categories of search content when clustering the rewritten search content each time can be the same or different; that is, the value of M can be different when clustering the rewritten search content each time.

[0205] When the second rewritten search result is retrieved from the N rewritten search results (i.e., i = 2), the M-class search results corresponding to this second rewritten search result will only include the search results created from the first rewritten search result. This class of search results now only includes the first rewritten search result. Similarly, the principle is the same when clustering subsequent rewritten search results. Clustering can be done using previously clustered search results, such as clustering the currently retrieved rewritten search result into a specific search result from the previously clustered search results, or creating a new search result based on the currently retrieved rewritten search result.

[0206] The above process can be used to cluster N rewritten search results to obtain at least one category of search results. The specific number of categories of search results can be determined according to the actual application scenario.

[0207] Step S203: Based on the search content within the target time slice of each type of search content, create a search content set corresponding to each type of search content of at least one type of search content.

[0208] Optionally, the computer device can create a search content set corresponding to each of the at least one type of search content within the target time slice, based on the search content within each of the at least one type of search content. Each type of search content corresponds to one search content set, and each search content set contains the search content within that type of search content whose search time falls within the target time slice. The search content within a search content set has content similarity (indicating that the search intent is the same or similar).

[0209] Therefore, explosive trending events (such as the aforementioned search outbreak events) can be discovered through the various search content sets obtained. The principle of discovering search outbreak events through each search content set is the same and independent, as detailed above. Figure 3 The process of mining search outbreak events described in the corresponding embodiment.

[0210] In this application, by rewriting and then clustering the various search contents to be processed, search contents with the same or similar intent can be processed uniformly. That is, by mining search outbreak events related to search contents by category, the accuracy of the mined search outbreak events can be improved. In addition, this application constructs a set of search contents by using the same time slices (such as target time slices) in different time periods, and then mines related search outbreak events together by the search contents within the time slice. This can avoid the problem of inconsistent and large differences in user search behavior in different time slices of a time period (such as the problem that the search volume is relatively large from 9:00 to 10:00 am, but relatively small from 11:00 pm to 12:00 am), and can also improve the accuracy of the mined search outbreak events.

[0211] Please see Figure 10 , Figure 10 This is a schematic diagram of an event mining scenario provided in this application. For example... Figure 10 As shown, s1: The computer device can acquire user search data. Here, user search data refers to the user's search content. User search data includes active search content (such as content entered by the user) and passive search content (such as terms clicked by the user on the interface). The user search data here can refer to the aforementioned original search content.

[0212] s2: The computer device can perform structural filtering on the acquired user search data (which can be denoted as a query, with one search term representing one query). This process can be referred to the above process of structural filtering on several original search terms. Next, s3: The computer device can perform semantic filtering on the structurally filtered search data. This process can be referred to the above process of semantic filtering on several structurally filtered search terms.

[0213] s4: Next, the computer can aggregate (i.e., cluster) the queries after semantic filtering. This process is similar to the clustering process for the N search terms described above. s5: The computer can calculate the QCV value (search burst coefficient) using the clustered query results through a sliding window (the size of a sliding window can be a target time slice; the sliding window can be continuously slid over time, and the target time slice will be continuously updated). s6: If the QCV value meets the burst coefficient criterion, the clustered rewritten search terms can be associated with corresponding resources. These resources can be at least one search result associated with the search term set mentioned above. s7: By associating resources, search burst events can be identified, and event titles for these identified events can be generated. s8: The computer can push the generated event titles to the front end, causing the front end to output the event (e.g., displaying the event title on the trending search list).

[0214] This application obtains a search content set corresponding to a target time segment; the search content set contains search content within multiple time segments of the target time segment, and the search content in the search content set has content similarity; it obtains the frequency fluctuation parameters of the search frequency of the search content in the search content set within the target time segments of multiple time segments; it determines the search burst coefficient of the search content in the search content set within the target time segments of multiple time segments based on the frequency fluctuation parameters; if the search burst coefficient meets the burst coefficient standard, it obtains the search burst events associated with the search content set within the target time segment. Therefore, the method proposed in this application can synchronously process search content with content similarity in the search content set, and obtain the search burst coefficient of the search content in the target time segment by using the frequency fluctuation coefficient of the search frequency of the search content in the same time segment (such as the target time segment) of each time segment. This takes into account the search burstability (such as volatility) of the search content, and thus obtains the search burst events through the search burstability of the search content, ensuring that the obtained search burst events also have search burstability, thus making the obtained search burst events more accurate.

[0215] Please see Figure 11 , Figure 11 This is a schematic diagram of the structure of a data processing apparatus provided in this application. The data processing apparatus can be a computer program (including program code) running on a computer device; for example, the data processing apparatus is application software. This data processing apparatus can be used to execute the corresponding steps in the methods provided in the embodiments of this application. Figure 11 As shown, the data processing device 1 may include: a set acquisition module 11, a parameter acquisition module 12, a coefficient determination module 13, and an event acquisition module 14;

[0216] The set acquisition module 11 is used to acquire the search content set corresponding to the target time segment; the search content set contains search content within the target time segment of multiple time periods, and the search content in the search content set has content similarity;

[0217] The parameter acquisition module 12 is used to acquire the frequency fluctuation parameters of the search frequency of the search content in the search content set within the target time segment of multiple time periods;

[0218] The coefficient determination module 13 is used to determine the search burst coefficient of the search content in the search content set within the target time slices of multiple time periods based on the frequency fluctuation parameter.

[0219] The event acquisition module 14 is used to acquire search burst events associated with the search content set within the target time slice if the search burst coefficient meets the burst coefficient standard.

[0220] Optionally, the methods by which the set acquisition module 11 acquires the search content set corresponding to the target time slice include:

[0221] Retrieve N search results to be processed; N search results contain search results from multiple time periods, where N is a positive integer;

[0222] Cluster the N search results for similar content to obtain at least one category of search results;

[0223] Based on the search content within the target time slice of each search content of at least one type of search content, create a search content set corresponding to each type of search content of at least one type of search content.

[0224] Optionally, the set acquisition module 11 performs clustering processing on N search results to obtain at least one category of search results, including:

[0225] Perform word segmentation on each of the N search terms to obtain the entity words contained in each of the N search terms;

[0226] Normalize and rewrite the entity segments contained in each of the N search terms to obtain the rewritten search terms corresponding to each of the N search terms.

[0227] For each rewritten search result, perform clustering of similar content to obtain at least one category of search results.

[0228] Optionally, the set acquisition module 11 performs clustering processing on similar content for each rewritten search content to obtain at least one category of search content, including:

[0229] Generate content embedding features for each rewritten search result;

[0230] Based on the content embedding features of each rewritten search content, clustering of similar content is performed on each rewritten search content to obtain at least one class of search content; any one of the at least one class of search content contains at least one rewritten search content.

[0231] Optionally, the number of rewritten search terms corresponding to N search terms is N, and any one of the N rewritten search terms is represented as the i-th rewritten search term, where i is a positive integer and i is less than or equal to N;

[0232] The collection acquisition module 11, based on the content embedding features of each rewritten search content, performs clustering processing on similar content for each rewritten search content to obtain at least one class of search content, including:

[0233] Get the M classes of search content clustered before clustering the i-th rewritten search content; each of the M classes of search content contains at least one rewritten search content, and M is a positive integer;

[0234] The content embedding features of the rewritten search content contained in each of the M types of search content are averaged to obtain the average embedding features of the rewritten search content contained in each of the M types of search content.

[0235] Based on the content embedding features of the i-th rewritten search content and the average embedding features corresponding to each of the M types of search content, the i-th rewritten search content is clustered to obtain at least one type of search content.

[0236] Optionally, the set acquisition module 11 performs clustering processing on the i-th rewritten search content based on the content embedding features of the i-th rewritten search content and the average embedding features corresponding to each of the M types of search content, to obtain at least one type of search content, including:

[0237] Obtain the feature distance between the content embedding feature of the i-th rewritten search content and the average embedding feature of each search content in the M-class search content. This feature distance is used to characterize the content similarity between the i-th rewritten search content and the rewritten search content contained in each search content in the M-class search content.

[0238] If there is a target class search content in the M-class search content with a corresponding feature distance less than or equal to the distance threshold, then the i-th rewritten search content is clustered into the target class search content to obtain the clustered target class search content.

[0239] If the target class search content does not exist in the search content of class M, then create a unique class search content based on the i-th rewritten search content;

[0240] Based on the search content of category M, excluding the target category search content after clustering, the target category search content after clustering, or the unique category search content, determine at least one type of search content.

[0241] Optionally, the parameter acquisition module 12 may acquire the frequency fluctuation parameters of the search frequency of search content in the search content set within multiple target time segments in the search content set in the following ways:

[0242] Obtain the search frequency of each search item in the search content set within the target time segment of each time period;

[0243] Based on the search frequency corresponding to each time period, calculate the frequency standard deviation of the search frequency of the search content in the search content set within the target time segment of multiple time periods.

[0244] The frequency standard deviation is defined as the frequency fluctuation parameter.

[0245] Optionally, any one of multiple time periods can be represented as the target time period; the parameter acquisition module 12 acquires the search frequency of the search content in the search content set within the target time segment of each time period in the following ways:

[0246] The search content that falls within the target time segment of the target time period is selected as the search content to be calculated.

[0247] Obtain the search type of each search item to be calculated, and determine the search frequency weight corresponding to each search item to be calculated based on the search type of each search item to be calculated.

[0248] Based on the search frequency weight corresponding to each search content to be calculated, the search frequency of the search content in the search content set within the target time segment of the target time period is determined.

[0249] Optionally, the parameter acquisition module 12 calculates the frequency standard deviation of the search frequency of the search content in the search content set within the target time slices of multiple time periods based on the search frequency corresponding to each time period. This includes methods such as:

[0250] The frequency decay weight for each time period is determined based on the chronological order between each time period.

[0251] The frequency standard deviation is determined based on the frequency decay weight and search frequency corresponding to each time period.

[0252] Optionally, the coefficient determination module 13 determines the search burst coefficient of the search content in the search content set within multiple time segments based on the frequency fluctuation parameter, including:

[0253] Based on the search frequency corresponding to each time period, determine the average search frequency of the search content in the search content set within the target time segments of multiple time periods.

[0254] The search burst coefficient is determined based on the average search frequency and frequency fluctuation parameters.

[0255] Optionally, if the search burst coefficient meets the burst coefficient standard, the event acquisition module 14 acquires the search burst events associated with the search content set within the target time slice in the following ways:

[0256] If the search burst coefficient is greater than or equal to the burst coefficient threshold, then the search burst coefficient is determined to meet the burst coefficient standard.

[0257] Retrieve at least one search result associated with the search content set;

[0258] A search outbreak event is determined based on at least one search result.

[0259] Optionally, the event acquisition module 14 may acquire at least one search result associated with the search content set in the following ways:

[0260] Select search content from the set of search results as a reference;

[0261] At least one search result is obtained based on the search content used as a reference.

[0262] Optionally, the event acquisition module 14 determines the search outbreak event based on at least one search result in the following ways:

[0263] Retrieve the topic text contained in each search result;

[0264] Each topic text is segmented into clauses to obtain multiple text clauses of the topic text contained in at least one search result;

[0265] Obtain the event evaluation index for each text clause, and determine the search burst event from multiple text clauses based on the event evaluation index for each text clause.

[0266] According to one embodiment of this application, Figure 3 The steps involved in the data processing method shown can be derived from... Figure 11 The data processing apparatus 1 shown is executed by each module. For example, Figure 3 Step S101 shown can be performed by Figure 11 The collection acquisition module 11 in the middle is used to execute, Figure 3 Step S102 shown can be performed by Figure 11 The parameter acquisition module 12 in the middle is used for execution; Figure 3 Step S103 shown can be performed by Figure 11 The coefficient determination module 13 in the middle is used for execution. Figure 3 Step S104 shown can be performed by Figure 11 The event acquisition module 14 in the middle is used to execute.

[0267] This application obtains a search content set corresponding to a target time slice; the search content set contains search content within multiple time slices of the target time period, and the search content in the search content set has content similarity; it obtains the frequency fluctuation parameters of the search frequency of the search content in the search content set within the target time slices of multiple time periods; it determines the search burst coefficient of the search content in the search content set within the target time slices of multiple time periods based on the frequency fluctuation parameters; if the search burst coefficient meets the burst coefficient standard, it obtains the search burst event associated with the search content set within the target time slice. Therefore, the device proposed in this application can synchronously process search content with content similarity in the search content set, and obtain the search burst coefficient of the search content in the target time slice by using the frequency fluctuation coefficient of the search frequency of the search content in the same time slice (such as the target time slice) in various time periods. This takes into account the search burstability (such as volatility) of the search content, and further obtains the search burst event through the search burstability of the search content, ensuring that the obtained search burst event also has search burstability, thus making the obtained search burst event more accurate.

[0268] According to one embodiment of this application, Figure 11 The modules in the data processing device 1 shown can be individually or entirely combined into one or more units, or some of the units can be further divided into multiple functionally smaller sub-units to achieve the same operation without affecting the technical effects of the embodiments of this application. The above modules are based on logical function division. In practical applications, the function of one module can also be implemented by multiple units, or the function of multiple modules can be implemented by one unit. In other embodiments of this application, the data processing device 1 may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0269] According to one embodiment of this application, a general-purpose computer device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM), can perform operations such as... Figure 3 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 11 The data processing apparatus 1 shown herein, and the data processing method for implementing the embodiments of this application, are described. The computer program described above may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned computing device via the computer-readable recording medium, and run therein.

[0270] Please see Figure 12 , Figure 12This is a schematic diagram of the structure of a computer device provided in this application. For example... Figure 12 As shown, the computer device 1000 may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 12 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0271] exist Figure 12 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0272] Obtain the search content set corresponding to the target time segment; the search content set contains search content within the target time segment across multiple time periods, and the search content within the search content set exhibits content similarity.

[0273] Obtain frequency fluctuation parameters of the search frequency of search content in the search content set within target time segments across multiple time periods;

[0274] The search burst coefficient of the search content in the search content set within the target time slices of multiple time periods is determined based on the frequency fluctuation parameter.

[0275] If the search burst coefficient meets the burst coefficient standard, then the search burst events associated with the search content set within the target time slice are obtained.

[0276] It should be understood that the computer device 1000 described in the embodiments of this application can execute the foregoing text. Figure 3 The data processing method described in the corresponding embodiments can also be executed as described above. Figure 11The description of the data processing apparatus 1 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.

[0277] Furthermore, it should be noted that this application also provides a computer-readable storage medium storing a computer program executed by the aforementioned data processing apparatus 1. The computer program includes program instructions, which, when executed by a processor, enable the execution of the aforementioned... Figure 3 The description of the data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the embodiments of the computer storage medium involved in this application, please refer to the description of the method embodiments of this application.

[0278] As an example, the above program instructions can be deployed and executed on a single computer device, or deployed and executed on multiple computer devices located in one location, or executed on multiple computer devices distributed across multiple locations and interconnected via a communication network. Multiple computer devices distributed across multiple locations and interconnected via a communication network can form a blockchain network.

[0279] The aforementioned computer-readable storage medium can be an internal storage unit of the data processing apparatus or computer device provided in any of the foregoing embodiments, such as a hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0280] This application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned... Figure 3The data processing method described in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.

[0281] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0282] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0283] The methods and related apparatuses provided in this application are described with reference to the method flowcharts and / or structural diagrams provided in this application. Specifically, each block of the method flowchart and / or structural diagram, as well as combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to create a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the process. Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 A process or multiple processes and / or structures illustrate the steps of the functions specified in one or more boxes.

[0284] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the search content set corresponding to the target time segment; the search content set contains search content within the target time segment in multiple time periods, and the search content in the search content set has content similarity; a time period includes multiple time intervals, and the target time segment refers to any time interval in a time period; Obtain the frequency fluctuation parameter of the search frequency of the search content in the search content set within the target time segment of the multiple time periods; The search burst coefficient of the search content in the search content set within the target time slice of the multiple time periods is determined based on the frequency fluctuation parameter. The search burst coefficient is used to measure whether the search content in the search content set reaches a burst state within the target time slice of the multiple time periods. If the search burst coefficient meets the burst coefficient standard, then the search burst events associated with the search content set within the target time slice are obtained.

2. The method according to claim 1, characterized in that, The process of obtaining the search content set corresponding to the target time slice includes: Obtain N search results to be processed; the N search results include search results to be processed within the multiple time periods, where N is a positive integer; Cluster the N search results for similar content to obtain at least one category of search results; Based on the search content within the target time slice of each of the at least one type of search content, create a set of search content corresponding to each of the at least one type of search content.

3. The method according to claim 2, characterized in that, The clustering of similar content among the N search results yields at least one category of search content, including: Each of the N search terms is segmented into words to obtain the entity words contained in each of the N search terms. The entity word segments contained in the N search contents are normalized and rewritten to obtain the rewritten search contents corresponding to the N search contents respectively; For each rewritten search content, perform clustering of similar content to obtain at least one category of search content.

4. The method according to claim 3, characterized in that, The step of clustering similar content for each rewritten search content to obtain at least one category of search content includes: Generate the content embedding features for each of the rewritten search contents; Based on the content embedding features of each rewritten search content, clustering of similar content is performed on each rewritten search content to obtain at least one class of search content; any one of the at least one class of search content contains at least one rewritten search content.

5. The method according to claim 4, characterized in that, The number of rewritten search contents corresponding to the N search contents is N. Any one of the N rewritten search contents is represented as the i-th rewritten search content, where i is a positive integer and i is less than or equal to N. Based on the content embedding features of each rewritten search content, clustering of similar content is performed on each rewritten search content to obtain at least one class of search content, including: Obtain the M classes of search content clustered before clustering the i-th rewritten search content; each of the M classes of search content contains at least one rewritten search content, and M is a positive integer; The content embedding features of the rewritten search content contained in each of the M types of search content are averaged to obtain the average embedding features of the rewritten search content contained in each of the M types of search content. Based on the content embedding features of the i-th rewritten search content and the average embedding features corresponding to each of the M types of search content, the i-th rewritten search content is clustered to obtain the at least one type of search content.

6. The method according to claim 5, characterized in that, The embedding features are based on the content embedding features of the i-th rewritten search content and the average embedding features corresponding to each of the M types of search content. Clustering is performed on the i-th rewritten search content to obtain the at least one class of search content, including: The feature distance between the content embedding feature of the i-th rewritten search content and the average embedding feature of each search content in the M-class search content is obtained. The feature distance is used to characterize the content similarity between the i-th rewritten search content and the rewritten search content contained in each search content in the M-class search content. If there is a target class search content in the M-class search content with a corresponding feature distance less than or equal to the distance threshold, then the i-th rewritten search content is clustered into the target class search content to obtain the clustered target class search content. If the target category search content does not exist in the M category search content, then a unique category search content is created based on the i-th rewritten search content; The at least one type of search content is determined based on the M types of search content excluding the clustered target class search content, the clustered target class search content, or the unique class search content.

7. The method according to claim 1, characterized in that, The step of obtaining the frequency fluctuation parameter of the search frequency of the search content in the search content set within the target time slice of the multiple time periods includes: Obtain the search frequency of each search item in the search content set within the target time segment of each time period; Based on the search frequency corresponding to each time period, calculate the frequency standard deviation of the search frequency of the search content in the target time segment of the multiple time periods; The frequency standard deviation is defined as the frequency fluctuation parameter.

8. The method according to claim 7, characterized in that, Any one of the multiple time periods is represented as the target time period; obtaining the search frequency of the search content in the search content set within the target time segment of each time period includes: The search content whose search time falls within the target time segment of the target time period is taken as the search content to be calculated; Obtain the search type of each search content to be calculated, and determine the search frequency weight corresponding to each search content to be calculated based on the search type of each search content to be calculated. Based on the search frequency weight corresponding to each search content to be calculated, the search frequency of the search content in the search content set within the target time segment of the target time period is determined.

9. The method according to claim 7, characterized in that, The step of calculating the frequency standard deviation of the search frequency of the search content in the search content set within the target time slices of the multiple time periods, based on the search frequency corresponding to each time period, includes: The frequency decay weight corresponding to each time period is determined according to the chronological order between each time period. The frequency standard deviation is determined based on the frequency decay weight and search frequency corresponding to each time period.

10. The method according to claim 7, characterized in that, The step of determining the search burst coefficient of the search content in the search content set within the target time slice of the multiple time periods based on the frequency fluctuation parameter includes: Based on the search frequency corresponding to each time period, determine the average search frequency of the search content in the search content set within the target time segment of the multiple time periods; The search burst coefficient is determined based on the average search frequency and the frequency fluctuation parameter.

11. The method according to claim 1, characterized in that, If the search burst coefficient meets the burst coefficient standard, then the search burst events associated with the search content set within the target time slice are obtained, including: If the search burst coefficient is greater than or equal to the burst coefficient threshold, then the search burst coefficient is determined to meet the burst coefficient standard. Obtain at least one search result associated with the search content set; The search outbreak event is determined based on the at least one search result.

12. The method according to claim 11, characterized in that, Obtaining at least one search result associated with the search content set includes: Select search content as a reference from the set of search content; The at least one search result is obtained based on the search content used as a reference.

13. The method according to claim 11, characterized in that, Determining the search outbreak event based on the at least one search result includes: Retrieve the topic text contained in each search result; Each topic text is segmented into clauses to obtain multiple text clauses of the topic text contained in the at least one search result; Obtain the event evaluation index for each text clause, and determine the search burst event from the plurality of text clauses based on the event evaluation index for each text clause.

14. A data processing apparatus, characterized in that, The device includes: The set acquisition module is used to acquire the search content set corresponding to the target time segment; the search content set contains search content within the target time segment in multiple time periods, and the search content in the search content set has content similarity; a time period includes multiple time intervals, and the target time segment refers to any time interval in a time period; The parameter acquisition module is used to acquire the frequency fluctuation parameter of the search frequency of the search content in the search content set within the target time segment of the multiple time periods; The coefficient determination module is used to determine the search burst coefficient of the search content in the search content set within the target time slice of the multiple time periods based on the frequency fluctuation parameter. The search burst coefficient is used to measure whether the search content in the search content set reaches a burst state within the target time slice of the multiple time periods. The event acquisition module is used to acquire search burst events associated with the search content set within the target time slice if the search burst coefficient meets the burst coefficient standard.

15. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-13.

16. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method according to any one of claims 1-13.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as described in any one of claims 1-13.