Data processing method, device and equipment, readable storage medium and program product

By filtering the quality feedback and content representation information of candidate response data, efficient and accurate matching data pairs are generated, which solves the problem of low manual labeling efficiency in the graphic consistency detection model, and achieves efficient and accurate matching of data.

CN120296237APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410045234.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the supervision and training of the graphic consistency detection model relies on manual annotation, which is inefficient and poorly accurate, making it difficult to efficiently and accurately generate matching data pairs.

Method used

By obtaining multiple candidate response data corresponding to the target text fragment, filtering using quality feedback information and content representation information to generate matching data pairs, including efficient and accurate matching of content data and title information.

Benefits of technology

It improves the generation efficiency and accuracy of matching data pairs, reduces the dependence of manual annotations, and ensures the authenticity and consistency of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296237A_ABST
    Figure CN120296237A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, device and equipment, a readable storage medium and a program product, which can be applied to fields or scenes such as cloud technology, artificial intelligence, intelligent platforms, application software, large models and vehicle-mounted application, and the method comprises the following steps: acquiring a plurality of candidate response data corresponding to a target text fragment; the multiple pieces of candidate response data are determined by the search engine in response to a search request for the target text fragment; according to the quality feedback information of each piece of candidate response data, multiple pieces of coarse screening response data are determined from the multiple pieces of candidate response data; determining multiple pieces of fine screening response data from the multiple pieces of coarse screening response data according to the content representation information of each piece of coarse screening response data; generating a plurality of matched data pairs according to the content data and the title information of the plurality of fine screening response data; the content data and the title information included in each matching data pair are matched. Through the embodiment of the invention, the matched data pair can be efficiently and accurately generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a data processing method, a data processing apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] The graphic-text consistency detection model plays a very important role in the Internet. Through the graphic-text consistency detection model, the authenticity of the graphic-text information published on the website can be verified. In order to improve the detection accuracy of the graphic-text consistency detection model, it is usually necessary to use matching data such as mutually matching text and video to supervise and train the model. At present, the matching data pair is mainly generated by manually annotating the text corresponding to the video, but the efficiency of manual annotation is extremely low, and it is difficult to completely avoid subjective judgment, resulting in low annotation accuracy. Therefore, how to efficiently and accurately generate matching data pairs is an urgent problem to be solved at present. Summary of the Invention

[0003] The present application provides a data processing method, apparatus, device, readable storage medium, and program product, which can efficiently and accurately generate matching data pairs.

[0004] In a first aspect, the present application provides a data processing method, which includes:

[0005] Obtain a plurality of candidate response data corresponding to a target text segment; the plurality of candidate response data are determined by a search engine in response to a search request for the target text segment, and the candidate response data include content data and title information, and the content data includes one or more of images and videos;

[0006] Determine a plurality of roughly screened response data from the plurality of candidate response data according to the quality feedback information of each candidate response data; each quality feedback information is determined according to the feedback data of the historical search object corresponding to the candidate response data for data quality;

[0007] Determine a plurality of finely screened response data from the plurality of roughly screened response data according to the content characterization information of each roughly screened response data; each content characterization information is determined according to the content data and title information of the corresponding roughly screened response data;

[0008] Generate a plurality of matching data pairs according to the content data and title information of the plurality of finely screened response data; the content data and title information included in each matching data pair match each other.

[0009] In another aspect, the present application provides a data processing apparatus, which includes:

[0010] An acquisition module, configured to acquire a plurality of candidate response data corresponding to a target text segment; the plurality of candidate response data are determined by a search engine in response to a search request for the target text segment, and the candidate response data includes content data and title information, and the content data includes one or more of images and videos;

[0011] A rough screening module, configured to determine a plurality of roughly screened response data from the plurality of candidate response data according to the quality feedback information of each candidate response data; each quality feedback information is determined according to the feedback data of the historical search object corresponding to the candidate response data for the data quality;

[0012] A fine screening module, configured to determine a plurality of finely screened response data from the plurality of roughly screened response data according to the content characterization information of each roughly screened response data; each content characterization information is determined according to the content data and title information of the corresponding roughly screened response data;

[0013] A processing module, configured to generate a plurality of matching data pairs according to the content data and title information of the plurality of finely screened response data; each matching data pair includes matching content data and title information.

[0014] On the other hand

[0015] Correspondingly, the present application provides a computer device, including: a processor, a storage device, and a communication interface, the processor, the communication interface, and the storage device are connected to each other, wherein the storage device stores executable program code, and the processor is configured to call the executable program code to implement the above data processing method.

[0016] Correspondingly, the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and the program instructions are executed by a processor to implement the data processing method as described above.

[0017] Correspondingly, the present application provides a computer program product, the computer program product includes a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the data processing method as described above.

[0018] The embodiments of the present application obtain multiple candidate response data corresponding to a target text segment. Each candidate response data includes content data and title information. Since the multiple candidate response data are determined by a search engine in response to a search request for the target text segment, this ensures the authenticity of the candidate response data. Then, through the dimension of quality feedback information, multiple roughly screened response data are initially screened out from the multiple candidate response data. Then, through the dimension of content representation information, multiple finely screened response data are further screened out from the multiple roughly screened response data. This ensures the matching degree between the content data and the title information of the finely screened response data obtained by screening. Finally, a matching data pair corresponding to each finely screened response data is generated. The content data and title information included in each matching data pair are highly matched. Compared with generating a matching data pair by manually annotating the text corresponding to the content data, the above method can generate matching data pairs more efficiently and accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a schematic structural diagram of a data processing system provided by an exemplary embodiment of the present application;

[0021] Figure 2 is a schematic flowchart of a data processing method provided by an exemplary embodiment of the present application;

[0022] Figure 3 is a schematic flowchart of another data processing method provided by an exemplary embodiment of the present application;

[0023] Figure 4A is a schematic diagram of title information and a display page provided by an exemplary embodiment of the present application;

[0024] Figure 4B is a processing flowchart of a graphic-text consistency detection model provided by an exemplary embodiment of the present application;

[0025] Figure 5 is a schematic structural diagram of a data processing device provided by an exemplary embodiment of the present application;

[0026] Figure 6 is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0028] The embodiments of the present application can be applied to various fields or scenarios such as cloud computing, cloud Internet of Things, cloud gaming, artificial intelligence, vehicle-mounted scenarios, intelligent transportation, and assisted driving. Several typical fields or scenarios will be introduced below.

[0029] Cloud computing refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources in a demand-driven and easily scalable manner through the network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services in a demand-driven and easily scalable manner through the network. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balance. With the development of the Internet, real-time data streams, and the diversification of connected devices, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from the previous parallel distributed computing, the emergence of cloud computing will, in concept, drive a revolutionary change in the entire Internet model and enterprise management model. The present application can store data such as target text fragments and candidate response data on a cloud server, and when the above different data needs to be used, it can be directly obtained on the cloud server, greatly improving the data acquisition speed.

[0030] Cloud gaming, also known as gaming on demand, is an online gaming technology based on cloud computing technology. Cloud gaming technology enables thin clients with relatively limited graphics processing and data computing capabilities to run high-quality games. In the cloud gaming scenario, the game does not run on the player's game terminal but on the cloud server. The cloud server renders the game scene into a video and audio stream and transmits it to the player's game terminal through the network. The player's game terminal does not need to have powerful graphics computing and data processing capabilities, but only needs to have basic streaming media playback capabilities and the ability to obtain the player's input instructions and send them to the cloud server. This application is applied to cloud gaming. When there is a corresponding business requirement, it generates a matching data pair related to cloud gaming and uses the matching data pair to train a graphic-text consistency detection model to improve the detection accuracy of the graphic-text consistency detection model in cloud gaming services.

[0031] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0032] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operating / interactive systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The solution provided in the embodiments of this application involves machine learning, computer vision technology, etc. under artificial intelligence technology, which will be described below:

[0033] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstrations. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0034] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further perform graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. Specifically, the method proposed in this application performs keyword extraction processing such as OCR on the display page (such as the cover of a video) to obtain the keywords in the display page, so as to facilitate subsequent data fine-screening operations through the keywords in the display page and the keywords in the title information.

[0035] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. This application will be illustrated through the following embodiments.

[0036] Please refer to Figure 1, this figure is a schematic diagram of the architecture of a data processing system provided by an exemplary embodiment of the present application. The data processing system may specifically include a terminal device 101 and a server 102. Among them, the terminal device 101 and the server 102 are connected through a network, for example, connected through a local area network, a wide area network, a mobile Internet, etc. The operation object operates on the browser or client application of the terminal device 101 to access various data. The server 102 can respond to this operation and provide various services related to data access for the operation object.

[0037] The terminal device 101 is also referred to as a terminal, a user equipment (UE), an access terminal, a user unit, a mobile device, a user terminal, a wireless communication device, a user agent, or a user device. The terminal device may be a smart home appliance, a handheld device with wireless communication function (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC)), a vehicle-mounted terminal, a smart voice interaction device, a wearable device, or other intelligent devices, etc., but is not limited thereto.

[0038] The server 102 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0039] In a possible implementation manner, the server 102 may obtain multiple candidate response data (such as multiple query results) corresponding to a target text segment (such as a certain search keyword) from each terminal device 101; the server 102 determines multiple roughly screened response data from the multiple candidate response data according to the quality feedback information of each candidate response data; each quality feedback information is determined according to the feedback data of the historical search object corresponding to the candidate response data for data quality (such as whether the object is clicked, the browsing duration of the object, etc.); the server 102 determines multiple finely screened response data from the multiple roughly screened response data according to the content representation information of each roughly screened response data; the server 102 finally generates multiple matching data pairs according to the content data and title information of the multiple finely screened response data; each matching data pair includes matching content data and title information.

[0040] Among them, the multiple candidate response data may be different objects of a search engine (such as Figure 1The objects 1, object 2, object n, etc. in [the system] are determined in response to a search request for a target text segment. The candidate response data includes content data and title information. For example, if multiple objects have all searched for "how to write a weekly report" on the terminal device 101 through a search engine, then "how to write a weekly report" is the target text segment. The images and videos displayed to each object by the search engine in response to the search operation are the candidate response data.

[0041] In a possible implementation, the architecture of the data processing system proposed in this application may further include a database, which can be used to store multiple candidate response data corresponding to different text segments respectively. For example, different objects perform search operations for different text segments on the terminal device 101 through a search engine. The terminal device 101 loads the search records corresponding to the search operations into the log data and stores the log data in the database (such as a log database) for subsequent analysis and processing. These data can be recorded in different database tables in the database. For example, the database can be a database provided in the server, that is, a database built-in or self-owned by the server; the database can also be an external database connected to the server, such as a cloud database (that is, a database deployed in the cloud), which can be specifically deployed based on any one of private cloud, public cloud, hybrid cloud, edge cloud, etc., so that the functions emphasized by the cloud database are different.

[0042] In a possible implementation, the terminal device 101 can store the log data in the database in the terminal device 101. Then, the server 102 can obtain multiple candidate response data corresponding to the target text segment (such as a certain search keyword) from the databases of each terminal device 101. The terminal device 101 can also store the log data in the database in the server 102. Then, the server 102 can obtain multiple candidate response data corresponding to the target text segment (such as a certain search keyword) from the database in the server 102.

[0043] It can be understood that the schematic diagram of the architecture of the system described in the embodiments of this application is for more clearly explaining the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided by the embodiments of this application. For example, the data processing method provided by the embodiments of this application can be executed not only by the server 102, but also by other servers or server clusters different from the server 102 and capable of communicating with the terminal device 101 and / or the server 102. As is known to those of ordinary skill in the art, Figure 1The number of terminal devices and servers in [it] is merely illustrative. According to the requirements of business implementation, terminal devices and servers with any number can be configured. Moreover, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems. In the subsequent embodiments, the above-mentioned terminal device 101 will be referred to as a terminal device, and the above-mentioned server 102 will be referred to as a server, which will not be elaborated in the subsequent embodiments.

[0044] Please refer to Figure 2 , which is a schematic flowchart of a data processing method provided by an exemplary embodiment of this application. Taking the application of this method to a server as an example for illustration, this method may include the following steps:

[0045] S201. Obtain multiple candidate response data corresponding to a target text segment; the multiple candidate response data are determined by a search engine in response to a search request for the target text segment, and the candidate response data includes content data and title information.

[0046] In a possible implementation manner, the content data may include one or more of images and videos.

[0047] In a possible implementation manner, the target text segment refers to the text segment input by a search object in a search engine for performing a data search task. For example, when the search object inputs "how to write a weekly report" in the search engine for data search, then "how to write a weekly report" is the target text segment. The multiple candidate response data refer to the data determined by the search engine in response to the search request of the search object for the target text segment. For example, the query results obtained by the search engine in response to the search request of "how to write a weekly report" are the candidate response data, such as the candidate response data may be video tutorials, text tutorials, note tutorials, etc. related to "how to write a weekly report".

[0048] In a possible implementation manner, the candidate response data may include content data and title information. The content data may include one of text, image, video, and audio. For example, the candidate response data may include the video content corresponding to a video tutorial and the title of the video content. Another example is that the candidate response data may include the text content corresponding to a text tutorial and the title of the text content. The content data may also include multiple of text, image, video, and audio. For example, the candidate response data may include the note content corresponding to a note tutorial and the title of the note content, and the note content includes text, images, etc. For the sake of easy understanding, in the subsequent embodiments, the candidate response data will be taken as an example of video data for illustration, that is, the content data is video content and the title information is the video title, which will not be elaborated in the subsequent embodiments.

[0049] In a possible implementation, the above-mentioned multiple candidate response data may refer to those determined by a search engine in response to search requests of multiple search objects for a target text segment at different times (or at the same time). For example, search object 1 searches for "how to write a weekly report" at times T1 and T2, search object 2 searches for "how to write a weekly report" at times T1 and T3, and search object 3 searches for "how to write a weekly report" at time T4. Since the search times are different, the candidate response data determined by the search engine at different times may be different or the same. The search engine finally determines the above-mentioned multiple candidate response data based on the candidate response data corresponding to each of the multiple search objects. For example, the candidate response data corresponding to each of the multiple search objects is integrated (the integration process may include data deduplication) to obtain multiple candidate response data.

[0050] It should be noted that the above-mentioned multiple candidate response data may also refer to those determined by a search engine in response to search requests of the same search object for a target text segment at different times. For example, search object 1 searches for "how to write a weekly report" at times T1, T2, and T3. The search engine finally determines the above-mentioned multiple candidate response data based on the candidate response data corresponding to search object 1. For example, the candidate response data corresponding to each of the multiple search objects is integrated (the integration process may include data deduplication) to obtain multiple candidate response data.

[0051] In a possible implementation, the server can obtain the above-mentioned multiple candidate response data from the log data corresponding to the search engine. For example, the server uses the API interface provided by the search engine to obtain search records according to the search keyword (such as the target text segment) and within a limited time range. The server can process and analyze the search records to obtain multiple candidate response data.

[0052] In a possible implementation, the form of the target text segment can be: keyword, key phrase, key sentence, and a combination of one or more of the above forms. For example, the target text segment can be "eat", "delicious food", "recommendations for delicious food near universities", "university food", "university food recommendations", etc.

[0053] In the embodiments of the present application, the server batch-obtains candidate response data through the target text segment in the search engine, which can improve the data acquisition efficiency and ensure the comprehensiveness of the obtained candidate response data. The server can subsequently screen the candidate response data to gradually determine the data whose content data matches the title information from the candidate response data. Since the candidate response data is determined by the search engine in response to the search request for the target text segment from the index library, the authenticity of the candidate response data can be guaranteed to avoid data forgery.

[0054] S202. Determine multiple roughly screened response data from multiple candidate response data according to the quality feedback information of each candidate response data.

[0055] In a possible implementation, each quality feedback information is determined according to the feedback data on data quality of the historical search objects corresponding to the candidate response data. That is to say, the quality feedback information of each candidate response data is determined according to the feedback data on data quality of the historical search objects corresponding to the candidate response data, and the feedback data on data quality is the data fed back by the historical search objects and can reflect the quality of the candidate response data. For example, the feedback data on data quality can refer to the browsing duration of the historical search objects for the candidate response data, whether to click to browse, whether to like, whether to collect, whether to forward, whether to comment, whether to recommend, etc.

[0056] In a possible implementation, the number of historical search objects can be one or more. By performing statistical, analytical, integrative, computational and other processing on the feedback data on data quality of each historical search object, the quality feedback information of each candidate response data can be obtained. For example, the quality feedback information of each candidate response data can include one or more of the average browsing duration, click-through rate, like count, collection count, forward count, comment count, and recommendation count of the candidate response data.

[0057] In a possible implementation, the server can judge the quality level of each candidate response data through the quality feedback information of each candidate response data. For example, if the average browsing duration and click-through rate of a certain candidate response data are relatively high, it means that the historical search objects are relatively satisfied with the candidate response data, which can indicate that the data quality of the candidate response data is relatively high. While if the average browsing duration and click-through rate of a certain candidate response data are relatively low, it means that the historical search objects are not very satisfied with the candidate response data, which can indicate that the data quality of the candidate response data is relatively low. And the relatively high data quality can be reflected in that the content data and title information of the candidate response data are relatively matched.

[0058] In the embodiment of the present application, the server determines multiple roughly screened response data from multiple candidate response data according to the quality feedback information of each candidate response data, realizing the preliminary filtering of the candidate response data with relatively low quality from the dimension of the feedback information of the historical search objects. The candidate response data with relatively low quality can be considered as those with a relatively low matching degree between the content data and the title information of the candidate response data, such as some "clickbait" videos. And the remaining candidate response data has relatively high quality and good user feedback, which to a certain extent shows that the content data and title information of the roughly screened response data obtained by filtering are relatively matched. The server can further screen the roughly screened response data in the subsequent process to obtain the finely screened response data.

[0059] S203. Determine multiple refined screening response data from multiple rough screening response data according to the content characterization information of each rough screening response data.

[0060] In a possible implementation, each content characterization information is determined according to the content data and title information of the corresponding rough screening response data. The server can extract the content characterization information from the content data and title information of each rough screening response data. The content characterization information refers to the characterization information of the specific data content of the rough screening response data, and through the content characterization information, the data characteristics of the rough screening response data in different dimensions can be characterized. For example, the content characterization information may include: the video title of the rough screening response data, the video tag (i.e., the attribution category of the video in the search engine), the video description, the video duration, the video cover frame, etc.

[0061] In a possible implementation, the content characterization information may correspond to two features, namely, the content feature corresponding to the content data and the title feature corresponding to the title information. The server can analyze the two features to determine whether the content feature matches the title feature. When the content feature matches the title feature, the corresponding rough screening response data is determined as the refined screening response data. Exemplarily, the server can determine the attribution category of the content data in the search engine according to the content feature, and determine the attribution category of the title information in the search engine according to the title feature. When the two attribution categories are the same category, it indicates that the content feature matches the title feature.

[0062] In the embodiment of the present application, the server determines multiple refined screening response data from multiple rough screening response data according to the content characterization information of each rough screening response data, realizing further filtering of the obtained rough screening response data from the dimension of the specific data content, so as to ensure that the content data and the title information of the filtered refined screening response data match each other.

[0063] S204. Generate multiple matching data pairs according to the content data and title information of the multiple refined screening response data; each matching data pair includes matching content data and title information.

[0064] In a possible implementation, each refined screening response data includes a content data and a title information. The server can combine the content data and the title information to obtain the matching data pair corresponding to the refined screening response data. Then, through the above method, the matching data pairs corresponding to each refined screening response data can be obtained, that is, multiple matching data pairs are obtained. The content data and the title information included in each matching data pair match each other. For example, the title information in the matching data pair is "How to write a weekly report", and the corresponding content data is a video tutorial related to "How to write a weekly report".

[0065] In the embodiments of the present application, the server generates a corresponding matching data pair for each refined screening response data. Since the refined screening response data is obtained by gradually screening a plurality of candidate response data corresponding to the target text segment according to the quality feedback information and the content representation information, this ensures a high degree of matching between the content data and the title information in the matching data pair. Compared with generating the matching data pair by manually annotating the text corresponding to the content data, the method provided by the embodiments of the present application can improve the accuracy and generation efficiency of the matching data pair.

[0066] Please refer to Figure 3 , which is a schematic flowchart of another data processing method provided by an exemplary embodiment of the present application. Taking the application of this method to a server as an example, the method may include the following steps:

[0067] S301. Obtain a plurality of candidate response data corresponding to the target text segment; the plurality of candidate response data are determined by the search engine in response to the search request for the target text segment, and the candidate response data includes content data and title information.

[0068] In a possible implementation manner, the search engine may store the search records of different historical search objects for the text segment to be searched at each historical moment into the log data, and the text segment to be searched includes the target text segment. In one case, the server may search for the target text segment through the search engine at the current moment to obtain the response data (i.e., the query result) of the search engine, and the server determines a plurality of candidate response data according to the above response data, such as taking the response data as the plurality of candidate response data, or performing data preprocessing (such as data deduplication, etc.) on the response data to obtain a plurality of candidate response data.

[0069] In another case, the server may obtain the search records of each historical search object for the target text segment at the historical moment from the log data, and the server determines a plurality of candidate response data according to the above search records, such as integrating the search records of each historical search object for the target text segment at the historical moment and taking them as the plurality of candidate response data, or performing data preprocessing (such as data deduplication, etc.) on the search records of each historical search object for the target text segment at the historical moment to obtain a plurality of candidate response data. Through the above method, the flexibility of obtaining candidate response data is improved.

[0070] In a possible implementation manner, the server may limit the time range, and then obtain a plurality of candidate response data corresponding to the target text segment within the time range. For example, the server may determine the video query results with the upload time within the time range among the multiple video query results obtained by searching for "singing competition" through the search engine at the current moment as the plurality of candidate response data, so as to ensure the timeliness of the candidate response data.

[0071] The method in step S202 will be introduced in detail through steps S302 - S304 as follows:

[0072] S302. Obtain the quality feedback information of the target candidate response data; the target candidate response data is any one of multiple candidate response data.

[0073] S303. Detect the data quality of the target candidate response data according to the quality feedback information of the target candidate response data to obtain a quality detection result.

[0074] In the above steps S302 - S303, taking the target candidate response data among multiple candidate response data as an example, the server detects the data quality of the target candidate response data according to the quality feedback information of the target candidate response data to obtain a quality detection result. The quality detection result is mainly obtained by analyzing the feedback information of the historical search objects of the target candidate response data.

[0075] In a possible implementation manner, the quality feedback information of the target candidate response data includes: the click count of the target candidate response data, and the average browsing duration. Based on this, the above step S303 can be implemented according to the following steps:

[0076] (1). Obtain the exposure count and total display duration of the target candidate response data.

[0077] The exposure count of the target candidate response data refers to: the total number of times the target candidate response data is displayed on the terminal devices of each historical search object. The total display duration of the target candidate response data refers to: the total duration during which the target candidate response data can be displayed on the terminal devices of historical search objects. Exemplarily, if the target candidate response data is a video, then the total display duration is the total duration of the video; if the target candidate response data is an image, then the total display duration is the preset display duration of the image.

[0078] (2). Determine the play rate according to the click count and exposure count, and determine the completion rate according to the average browsing duration and total display duration.

[0079] The click count of the target candidate response data refers to: the total number of times the target candidate response data is clicked after being displayed on the terminal devices of each historical search object. The average browsing duration of the target candidate response data refers to: the average duration for multiple historical search objects to browse the target candidate response data. The play rate and completion rate can help measure the attractiveness of the target candidate response data to historical search objects. By analyzing the play rate and completion rate of the target candidate response data, the quality of the data of the target candidate response data can be judged good or bad.

[0080] Exemplarily, the search text fragment is denoted as query, and the response data corresponding to the search text fragment is denoted as doc. There are multiple docs corresponding to one search text fragment, that is, one query corresponds to a doc list. A data pair pair can be constructed between a query and any doc in the doc list. The server can obtain the relevant information of the doc in each pair, such as the click count record, the exposure count record, the average browsing duration, and the total display duration. The click count is denoted as Click, the exposure count is denoted as Exposure, the average browsing duration is denoted as Review, and the total display duration is denoted as Duration. Then, each data pair can be denoted as:

[0081] pair: <Exposure, Click, Duration, Review>

[0082] The playback rate is denoted as Fun1(doc), and the complete playback rate is denoted as Fun2(doc). Then, the calculation formulas for the playback rate and the complete playback rate can be as follows:

[0083] Fun1(doc) = Click / Exposure

[0084] Fun2(doc) = Review / Duration

[0085] (3) When the playback rate is greater than or equal to the first threshold and the complete playback rate is greater than or equal to the second threshold, determine that the quality detection result of the target candidate response data passes the detection.

[0086] In a possible implementation manner, the server can determine the first threshold according to the average click probability (i.e., the average playback probability) of the multiple candidate response data corresponding to the target text fragment and the first hyperparameter. For example, the product of the average click probability and the first hyperparameter is used as the first threshold.

[0087] Exemplarily, the average click probability of the multiple candidate response data is denoted as avg_p, and the first hyperparameter is denoted as x_param. Then, the first threshold can be denoted as avg_p * x_param. x_param can take values in the range of (1, 2].

[0088] Exemplarily, the second threshold is denoted as y_param, and y_param can take values in the range of (0.5, 1].

[0089] When the server determines that Fun1(doc) ≥ avg_p * x_param and Fun2(doc) ≤ y_param, it determines that the quality inspection result of the target candidate response data passes the inspection. In this scenario, the server considers the target candidate response data to have a high playback rate and a completion rate exceeding half of the duration before it deems the target candidate response data as high-quality data, rather than low-quality data such as clickbait. This largely indicates that the content data and title information of the target candidate response data are matched.

[0090] In a possible implementation, the quality feedback information of the target candidate response data may include: the number of clicks, average browsing duration, number of comments, number of likes, number of collections, number of recommendations, number of forwards, etc. of the target candidate response data. Then, the server can detect the data quality of the target candidate response data based on one or more of the data in the quality feedback information to obtain the quality inspection result.

[0091] Exemplarily, the server can analyze the three indicators of the number of comments, number of likes, and number of collections. When the number of comments, number of likes, and number of collections meet the preset conditions, it determines that the quality inspection result of the target candidate response data passes the inspection. Among them, meeting the preset conditions may mean that one or more of the number of comments, number of likes, and number of collections meet the determination thresholds of the corresponding indicators. It should be noted that the specific implementation of the quality inspection can be flexibly adjusted according to the actual situation and will not be elaborated here.

[0092] S304. If the quality inspection result is passed, the target candidate response data is determined as the roughly screened response data, so as to determine multiple roughly screened response data from multiple candidate response data.

[0093] In the embodiments of the present application, the server determines the candidate response data with a passed quality inspection result among the multiple candidate response data as the roughly screened response data, so that a part of the data with relatively matching content data and title information can be screened out from a large number of candidate response data. This helps to reduce the amount of data for subsequent fine screening, improve the efficiency of fine screening, and thus improve the generation efficiency of matching data pairs.

[0094] The method in step S203 will be introduced in detail through steps S305 - S307 below:

[0095] S305. Obtain the content representation information of the target roughly screened response data; the target roughly screened response data is any one of the multiple roughly screened response data.

[0096] S306. Detect the data content of the target roughly screened response data according to the content representation information of the target roughly screened response data to obtain the content inspection result.

[0097] In the above steps S305 - S306, taking the target rough screening response data among multiple rough screening response data as an example, the server detects the content quality of the content representation information of the target rough screening response data according to the content representation information of the target rough screening response data, and obtains a content detection result, which is mainly obtained by analyzing the data content of the target rough screening response data.

[0098] In a possible implementation manner, the content representation information of the target rough screening response data includes: the display page of the content data of the target rough screening response data, and the title information. Based on this, the above step S306 can be implemented according to the following steps:

[0099] (1) Extract keywords from the title information to obtain M first keywords, and extract keywords from the display page to obtain N second keywords.

[0100] In a possible implementation manner, the following methods can be used to extract keywords from the title information (such as a piece of text): word segmentation, word frequency statistics, part-of-speech statistics, term frequency - inverse document frequency (TF-IDF), etc.

[0101] Exemplarily, the title information is denoted as Title_terms, and the server can perform word segmentation processing on Title_terms to obtain the following word string:

[0102] Title_terms = [t1, t2…, tm]

[0103] Among them, t1, t2, etc. are respectively multiple word fragments obtained through word segmentation processing, and here m = M.

[0104] In a possible implementation, the display page may refer to the video cover of content data (such as a video). Then, the following methods can be used to perform keyword extraction on the display page (such as a frame image in the video): optical character recognition (OCR) and word segmentation, image keyword extraction based on visual features, image label extraction based on deep learning, etc. Among them, image keyword extraction based on visual features may refer to using image processing technology to extract visual features such as the color, texture, and shape of the image, and then determining corresponding keywords based on these features. Image label extraction based on deep learning may refer to using a pre-trained deep learning model, such as a Convolutional Neural Network (CNN), to perform feature extraction and image label prediction on the image to obtain corresponding keywords. It should be noted that when the content data is a video, the above display page may be any one or more video frames in the video, which will not be elaborated here.

[0105] Exemplarily, the OCR recognition result of the display page is denoted as Ocr_trems. The server can perform word segmentation processing on Ocr_trems to obtain the following word strings:

[0106] Ocr_trems = [k1, k2…, kn]

[0107] Among them, k1, k2, etc. are respectively multiple word segments obtained through word segmentation processing, and here n = N.

[0108] (2) Determine K identical keywords among the M first keywords and the N second keywords.

[0109] Exemplarily, the server can perform cross-comparison at the word granularity on the two word strings Title_terms and Ocr_trems to obtain a comparison result. The calculation formula can be as follows:

[0110] Ratio_cross = same(Title_terms, Ocr_trems) / max(Title_terms, Ocr_trems)

[0111] Among them, the same(,) function is used to calculate the number of common words between two word strings, such as obtaining K identical keywords. The max(,) function is used to calculate the maximum length of two word strings. For example, the maximum length is P, and P is the maximum value of M and N.

[0112] In a possible implementation, the server can also perform cross-comparison at the character granularity on the two word strings Title_terms and Ocr_trems to obtain a comparison result, which will not be elaborated in the embodiments of this application.

[0113] (3) When the ratio of K to P is greater than or equal to the third threshold, determine that the content detection result of the target rough screening response data passes the detection; M, N, and K are non-negative integers, and P is the maximum value of M and N.

[0114] Exemplarily, the third threshold is denoted as z_params, such as 0.3. When the ratio of K to P is greater than or equal to the third threshold, that is, Ratio_cross ≥ z_params, determine that the content detection result of the target rough screening response data passes the detection; when the ratio of K to P is less than the third threshold, that is, Ratio_cross < z_params, determine that the content detection result of the target rough screening response data fails the detection. Among them, z_params can be flexibly set according to service data and experience, and the embodiments of the present application do not limit this.

[0115] Through the above steps (1)-(3), the server can avoid complex calculations on the target rough screening response data. By extracting and comparing keywords from the title information of the target rough screening response data and the display page, the server can quickly and accurately detect the data content, thereby improving the generation efficiency of the matching data pairs.

[0116] As Figure 4A shown, Figure 4A is a schematic diagram of title information and a display page provided by an exemplary embodiment of the present application. The server processes the title information to extract 8 first keywords, namely: challenge tournament, strongest, king, god level, amazing operation, dare, come, and fight, and processes the display page to extract 4 second keywords, namely: challenge tournament, strongest, king, and amazing operation. At this time, K is 4, P is 8, the ratio of K to P is 0.5, which is greater than the third threshold (such as 0.3). Then, it is considered that the content detection result of the target rough screening response data passes the detection.

[0117] S307. If the content detection result passes the detection, determine the target rough screening response data as the fine screening response data, so as to determine multiple fine screening response data from multiple rough screening response data.

[0118] In the embodiments of the present application, the server determines the rough screening response data with a passing quality detection result among multiple rough screening response data as the fine screening response data, so that a part of the content data that is relatively matched with the title information can be further screened out from a large number of rough screening response data. The calculation amount of the above method is small, which ensures the efficiency of the fine screening and thus ensures the generation efficiency of the matching data pairs.

[0119] S308. Generate multiple matching data pairs based on the content data and title information of multiple finely screened response data; the content data and title information included in each matching data pair match each other.

[0120] In the above steps S301 - S308, the server first obtains multiple candidate response data corresponding to the target text segment. Each candidate response data includes content data and title information. Since the multiple candidate response data are determined by the search engine in response to the search request for the target text segment, this ensures the authenticity of the candidate response data. Then, through the dimension of quality feedback information, multiple coarsely screened response data are preliminarily screened out from the multiple candidate response data, and then through the dimension of content representation information, multiple finely screened response data are further screened out from the multiple coarsely screened response data, which ensures the matching degree between the content data and title information of the finely screened response data obtained by screening. Finally, a matching data pair corresponding to each finely screened response data is generated. The content data and title information included in each matching data pair match highly. Compared with generating matching data pairs by manually annotating the text corresponding to the content data, the above method can generate matching data pairs more efficiently and accurately.

[0121] The data processing method provided by the embodiments of the present application can be applied in the field of data mining. Taking the video search scenario in a video search engine as an example, the query results of the video search engine include multiple groups of data, and each group of data consists of a video and a video title. However, there may be a problem of inconsistent graphics and text between the video and the video title, that is, the video content does not match the video title. For example, the video content of a certain group of data is related to game A, while the video title is related to game B. The reasons for the inconsistent graphics and text can include the following two points: The first point is that abnormal situations occur during the process of obtaining web page data through methods such as crawlers and parsing the web page data to obtain the video and the video title. The second point is that the user uploads a video title that does not match the video content to improve the video's attractiveness, that is, a "clickbait" video.

[0122] To address the above problems, a graphics and text consistency detection model can be used to detect the consistency between the video and the video title. As Figure 4B shown, Figure 4B is a processing flowchart of a graphics and text consistency detection model provided by an exemplary embodiment of the present application. By inputting the video (such as content data) and the title (such as title information) into the graphics and text consistency detection model for processing, the graphics and text consistency detection model will process the video and the title respectively through the encoding module, and then fuse the processing results corresponding to the video and the processing results corresponding to the title through the fusion module, and further can determine the detection result. The detection result is used to indicate whether the video and the title match or do not match.

[0123] The text-image consistency detection model is usually obtained based on model training to ensure the detection accuracy of the text-image consistency detection model. For example, the initial detection model is trained based on the supervised training method to obtain the text-image consistency detection model. However, the supervised training method requires a large amount of labeled data (such as text-image consistency data, and the text in the text-image consistency data can be used as the supervision data of the video) to support sufficient training, and how to construct a large amount of text-image consistency data is the key to improving the model detection accuracy. The data processing method proposed in the embodiments of the present application can be regarded as a text-image consistency data mining solution based on the log data of the search engine, which can efficiently and accurately generate text-image consistency data (i.e., matching data pairs), effectively avoid the influence of the subjectivity of the manual annotation method on the objectivity of the data, and reduce the construction cost of the supervision data.

[0124] The method for batch constructing text-image consistency data (i.e., matching data pairs) will be described below:

[0125] In a possible implementation manner, the server can perform the following steps:

[0126] (1) Determine the text segment set; the text segment set includes multiple text segments, and the target text segment is any one of the multiple text segments.

[0127] In the embodiments of the present application, the server needs to determine multiple text segments to form the text segment set, and these text segments are the basis for batch constructing text-image consistency data.

[0128] (2) Determine multiple matching data pairs corresponding to each text segment, and generate a training sample set according to the multiple matching data pairs corresponding to each text segment.

[0129] In the embodiments of the present application, for each text segment in the text segment set, the server can use the method in the foregoing embodiments to determine multiple matching data pairs corresponding to each text segment, and then integrate the multiple matching data pairs corresponding to all text segments to obtain a training sample set. The integration process may include data deduplication processing. Each training sample in the training sample set corresponds to a matching data pair, that is, the training sample includes content data (such as a video) and title information (such as a title), and the title information can be used as the annotation information of the content data.

[0130] (3) Use the training samples in the training sample set to train the initial detection model to obtain the text-image consistency detection model.

[0131] In the embodiments of the present application, the server can use the training samples in the training sample set to train the initial detection model, so as to obtain a graphic-text consistency detection model. Through the way of supervised training, the model can learn the correlation features between the video and the title in the matching data pair, so as to improve the prediction accuracy of the model.

[0132] It should be noted that the above training sample set can be regarded as a positive sample set, and the training samples in the positive sample set are graphic-text consistency data. The server can also obtain the filtered candidate response data from the multiple candidate response data corresponding to each text segment, determine the unmatched data pairs, and then generate a negative sample set according to the multiple unmatched data pairs corresponding to each text segment. The training samples in the negative sample set are graphic-text inconsistency data. The server can use the positive sample set and the negative sample set to train the initial detection model. The joint use of positive and negative samples can help the model learn the consistency features and inconsistency features between images and texts, improve the generalization ability and robustness of the model, and avoid overfitting and underfitting of the model.

[0133] Through the above steps (1)-(3), batch generation of graphic-text consistency data can be realized, the data construction efficiency can be improved, the data construction cost can be reduced, and the authenticity and accuracy of the graphic-text consistency data can be ensured.

[0134] In a possible implementation manner, the above determination of the text segment set can be implemented according to the following steps:

[0135] (1) Determine multiple historical search text segments within a target time range according to the log data of the search engine.

[0136] In a possible implementation manner, when the user performs a search operation on the terminal device through the search engine for a search text segment (such as a keyword), the terminal device will load the search record corresponding to the search operation into the log data and store the log data in a database (such as a log database) for subsequent analysis and processing. The server can obtain the log data of the search engine from the database, and then screen out multiple historical search text segments within a certain time range from the log data. For example, the server can count all the search text segments input by users on the search engine within half a year to obtain multiple historical search text segments.

[0137] Exemplarily, the server performs data mining based on the log data of a mature search engine. Through the queries (i.e., search text fragments) initiated by users recently, it collects the doc list corresponding to each query (the doc list includes multiple docs, and a doc is the response data corresponding to the search text fragment), thereby obtaining a large number of Q-D data pairs. Here, Q refers to a query, and D refers to the doc list corresponding to the query. At the same time, the server can collect the quality feedback information corresponding to each doc, such as the number of clicks, average browsing duration, number of likes, number of collections, number of forwards, number of comments, number of recommendations, etc. At the same time, it can also obtain the total display duration of the exposure times, etc.

[0138] The collection of queries is very important. The collection of queries needs to take into account various types of docs in the search engine to ensure the dimensional diversity of the text fragments in the generated text fragment set. In the embodiments of the present application, the text fragment set will be constructed from two aspects: the search popularity of the query and the search category of the query.

[0139] (2) Determine the first historical search text fragment from multiple historical search text fragments according to multiple search frequency intervals.

[0140] In the embodiments of the present application, the server can perform frequency statistics processing on multiple historical search text fragments from the dimension of search popularity, obtain the search frequencies corresponding to different text fragments respectively, then according to multiple search frequency intervals, determine a certain number of text fragments from each search frequency interval, and then integrate the text fragments determined from multiple search frequency intervals respectively to obtain the first historical search text fragment.

[0141] Perform data mining based on the search frequency of the query in the recent period. Here, the period can be one month, half a year, one year, etc. Different queries are divided into multiple frequency levels according to the search frequency, such as high frequency, medium frequency, and low frequency. Among them, different frequency levels of queries are constructed according to a preset ratio. For example, the ratio of high frequency, medium frequency, and low frequency is 4:4:2. It should be noted that the proportions of high frequency and medium frequency are relatively large here, mainly considering that the docs corresponding to high-frequency and medium-frequency queries basically cover most of the data in the index library of the search engine.

[0142] Exemplarily, the multiple search frequency ranges can be: [0, 1000], [1000, 2000], [2000, positive infinity]. The server can determine 4*F high-frequency queries in the range of [2000, positive infinity], 4*F medium-frequency queries in the range of [1000, 2000], and 2*F low-frequency queries in the range of [0, 1000], where F is a positive integer. Then, the queries determined in each range are integrated to obtain the first historical search text fragment. At this time, the number of the first historical search text fragments is 10*F.

[0143] Through the above method, the text fragments in the first historical search text fragment have better balance, avoiding over-reliance on high-frequency queries or ignoring low-frequency queries. Generating a corresponding training sample set based on the above first historical search text fragment and performing model training can improve the generalization ability of the model.

[0144] (3) Determine the second historical search text fragment from multiple historical search text fragments according to multiple search categories.

[0145] In the embodiment of the present application, the server can perform category statistics processing on multiple historical search text fragments from the dimension of search categories to obtain the search categories corresponding to different text fragments respectively. Then, multiple search categories are preset, and a certain number of text fragments are determined from each preset search category. Then, the text fragments determined from multiple search categories are integrated to obtain the second historical search text fragment.

[0146] Data mining is performed based on search categories. Here, the search categories can be roughly divided into sports, games, entertainment, finance, etc. The server needs to take into account the distribution of queries in each category for data extraction. In different types of search engines, the category types of queries are generally different. For example, in general-purpose search engines, the search demands for categories such as entertainment and finance are more abundant, while vertical search has more abundant search demands in specific fields, such as the game field. In the embodiment of the present application, queries are extracted according to search categories, which can effectively evaluate the data distribution in the sample data and make the data distribution closer to the real business data distribution.

[0147] Exemplarily, the server can extract a certain number of text fragments from each of the four search categories of sports, games, entertainment, and finance, and obtain the second historical search text fragment through integration. Among them, the extraction quantities corresponding to the four search categories can be the same or different. The extraction quantities corresponding to different search categories can be flexibly set according to the actual business situation, and the embodiment of the present application does not limit this.

[0148] Through the above method, it can be ensured that the second historical search text fragments cover different search categories, guaranteeing the comprehensiveness and diversity of the data, and making the text fragments in the second historical search text fragments have better balance, avoiding over-reliance on a certain search category or ignoring other categories. Generating corresponding training samples according to the above second historical search text fragments and performing model training can improve the generalization ability of the model. Improve the robustness and generalization ability of the model.

[0149] (4) Determine a text fragment set according to the first historical search text fragment and the second historical search text fragment.

[0150] In the embodiment of the present application, the server determines a text fragment set according to the first historical search text fragment and the second historical search text fragment, ensuring the comprehensiveness and richness of the text fragments in the text fragment set. Generating corresponding training samples according to the text fragment set and performing model training can improve the generalization ability of the model and further improve the model prediction accuracy.

[0151] In a possible implementation manner, the first historical search text fragment and the second historical search text fragment can be integrated to obtain a text fragment set. The server can also extract a certain number of text fragments from the first historical search text fragment, extract a certain number of text fragments from the second historical search text fragment, and then integrate the extracted text fragments to obtain a text fragment set.

[0152] In a possible implementation manner, the server can first determine multiple search categories, then respectively determine text fragments corresponding to multiple frequency levels from each search category, and then integrate all the determined text fragments to obtain a text fragment set. Exemplarily, taking the sports category among multiple search categories as an example, the server obtains multiple historical search text fragments under the sports category, and then extracts corresponding numbers of text fragments from the multiple historical search text fragments according to the ratio of high frequency: medium frequency: low frequency being 4:4:2.

[0153] Based on the above embodiments, the beneficial effects of the present application are as follows: The embodiment of the present application performs automatic mining of image-text consistency data based on the log data of the search engine, which can effectively avoid problems such as low efficiency, high cost, and subjective deviation caused by the manual annotation method. By establishing a text fragment set and performing rough screening and fine screening on the candidate response data of each text fragment in turn, the accuracy of the matching data pairs is improved and the generation is efficient. The above method can be applied to establish basic capabilities such as image-text consistency and question-text consistency in various search scenarios, such as establishing an image-text consistency detection model and a question-text consistency detection model, thereby helping to screen out low-quality videos and articles with inconsistent images and texts or inconsistent questions and texts.

[0154] Please refer to Figure 5, which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. Specifically, the data processing device may include:

[0155] An acquisition module 501, configured to acquire a plurality of candidate response data corresponding to a target text segment; the plurality of candidate response data are determined by a search engine in response to a search request for the target text segment, and the candidate response data includes content data and title information, and the content data includes one or more of images and videos;

[0156] A rough screening module 502, configured to determine a plurality of rough screening response data from the plurality of candidate response data according to the quality feedback information of each candidate response data; each quality feedback information is determined according to the feedback data of the historical search object corresponding to the candidate response data for data quality;

[0157] A fine screening module 503, configured to determine a plurality of fine screening response data from the plurality of rough screening response data according to the content characterization information of each rough screening response data; each content characterization information is determined according to the content data and title information of the corresponding rough screening response data;

[0158] A processing module 504, configured to generate a plurality of matching data pairs according to the content data and title information of the plurality of fine screening response data; each matching data pair includes matching content data and title information.

[0159] In a possible implementation manner, when the rough screening module 502 is configured to determine a plurality of rough screening response data from the plurality of candidate response data according to the quality feedback information of each candidate response data, it is specifically configured to:

[0160] Acquire the quality feedback information of the target candidate response data; the target candidate response data is any one of the plurality of candidate response data;

[0161] Detect the data quality of the target candidate response data according to the quality feedback information of the target candidate response data, and obtain a quality detection result;

[0162] If the quality detection result is a pass, determine the target candidate response data as the rough screening response data.

[0163] In a possible implementation manner, the quality feedback information of the target candidate response data includes: the click-through rate and the average browsing duration of the target candidate response data;

[0164] Among them, when the above-mentioned rough screening module 502 is used to detect the data quality of the above-mentioned target candidate response data according to the quality feedback information of the above-mentioned target candidate response data and obtain a quality detection result, it is specifically used for:

[0165] Obtain the exposure times and total display duration of the above-mentioned target candidate response data;

[0166] Determine the play rate according to the above-mentioned click times and exposure times, and determine the completion rate according to the above-mentioned average browsing duration and total display duration;

[0167] When the above-mentioned play rate is greater than or equal to the first threshold and the above-mentioned completion rate is greater than or equal to the second threshold, determine that the quality detection result of the above-mentioned target candidate response data is passed.

[0168] In a possible implementation manner, when the above-mentioned refined screening module 503 is used to determine multiple refined screening response data from the above-mentioned multiple rough screening response data according to the content characterization information of each above-mentioned rough screening response data, it is specifically used for:

[0169] Obtain the content characterization information of the target rough screening response data; the target rough screening response data is any one of the above-mentioned multiple rough screening response data;

[0170] Detect the data content of the above-mentioned target rough screening response data according to the content characterization information of the above-mentioned target rough screening response data, and obtain a content detection result;

[0171] If the above-mentioned content detection result is passed, determine the above-mentioned target rough screening response data as the refined screening response data.

[0172] In a possible implementation manner, the content characterization information of the above-mentioned target rough screening response data includes: the display page of the content data of the above-mentioned target rough screening response data, and the title information;

[0173] Among them, when the above-mentioned refined screening module 503 is used to detect the data content of the above-mentioned target rough screening response data according to the content characterization information of the above-mentioned target rough screening response data and obtain a content detection result, it is specifically used for:

[0174] Perform keyword extraction processing on the above-mentioned title information to obtain M first keywords, and perform keyword extraction processing on the above-mentioned display page to obtain N second keywords;

[0175] Determine K identical keywords among the above-mentioned M first keywords and the above-mentioned N second keywords;

[0176] When the ratio of K to P is greater than or equal to the third threshold, it is determined that the content detection result of the above-mentioned target rough screening response data passes the detection; M, N, and K are non-negative integers, and P is the maximum value of M and N.

[0177] In a possible implementation manner, the above-mentioned processing module 504 is further configured to:

[0178] Determine a text segment set; the above-mentioned text segment set includes multiple text segments, and the above-mentioned target text segment is any one of the above-mentioned multiple text segments;

[0179] Determine multiple matching data pairs corresponding to each of the above-mentioned text segments, and generate a training sample set according to the multiple matching data pairs corresponding to each of the above-mentioned text segments;

[0180] Use the training samples in the above-mentioned training sample set to train the initial detection model to obtain a graphic-text consistency detection model.

[0181] In a possible implementation manner, when the above-mentioned processing module 504 is used to determine the text segment set, it is specifically configured to:

[0182] Determine multiple historical search text segments within a target time range according to the log data of the above-mentioned search engine;

[0183] Determine the first historical search text segment from the above-mentioned multiple historical search text segments according to multiple search frequency intervals;

[0184] Determine the second historical search text segment from the above-mentioned multiple historical search text segments according to multiple search categories;

[0185] Determine the text segment set according to the above-mentioned first historical search text segment and the above-mentioned second historical search text segment.

[0186] It should be noted that the functions of the functional modules of the data processing device in the embodiments of the present application can be specifically implemented according to the methods in the above-mentioned method embodiments, and the specific implementation process can refer to the relevant descriptions of the above-mentioned method embodiments, which will not be elaborated here.

[0187] Please refer to Figure 6 , this figure is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 6 shown, the computer device in this embodiment may include: a processor 601, a storage device 602, and a communication interface 603. Data interaction can be performed between the above-mentioned processor 601, storage device 602, and communication interface 603.

[0188] The above storage device 602 may include a volatile memory, such as a random-access memory (RAM); the storage device 602 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the above storage device 602 may further include a combination of the above types of memories.

[0189] The above processor 601 may be a central processing unit (CPU). In one embodiment, the above processor 601 may also be a Graphics Processing Unit (GPU). The above processor 601 may also be a combination of a CPU and a GPU. In one embodiment, the above storage device 602 is used to store program instructions, and the above processor 601 may call the above program instructions to perform the following operations:

[0190] Obtain a plurality of candidate response data corresponding to the target text segment; the above plurality of candidate response data are determined by the search engine in response to a search request for the above target text segment, and the candidate response data includes content data and title information, and the content data includes one or more of images and videos;

[0191] Determine a plurality of roughly screened response data from the above plurality of candidate response data according to the quality feedback information of each above candidate response data; each above quality feedback information is determined according to the feedback data of the historical search object corresponding to the candidate response data for the data quality;

[0192] Determine a plurality of finely screened response data from the above plurality of roughly screened response data according to the content characterization information of each above roughly screened response data; each above content characterization information is determined according to the content data and title information of the corresponding roughly screened response data;

[0193] Generate a plurality of matching data pairs according to the content data and title information of the above plurality of finely screened response data; each above matching data pair includes matching content data and title information.

[0194] In a possible implementation manner, when the above processor 601 is used to determine a plurality of roughly screened response data from the above plurality of candidate response data according to the quality feedback information of each above candidate response data, it is specifically used for:

[0195] Obtain the quality feedback information of the target candidate response data; the above target candidate response data is any one of the above plurality of candidate response data;

[0196] Detect the data quality of the above-mentioned target candidate response data according to the quality feedback information of the above-mentioned target candidate response data, and obtain a quality detection result;

[0197] If the above quality detection result is passed, determine the above target candidate response data as the roughly screened response data.

[0198] In a possible implementation, the quality feedback information of the above-mentioned target candidate response data includes: the number of clicks and the average browsing duration of the above-mentioned target candidate response data;

[0199] Among them, when the above-mentioned processor 601 is used to detect the data quality of the above-mentioned target candidate response data according to the quality feedback information of the above-mentioned target candidate response data and obtain a quality detection result, it is specifically used for:

[0200] Obtain the exposure times and the total display duration of the above-mentioned target candidate response data;

[0201] Determine the playback rate according to the above number of clicks and the above exposure times, and determine the completion rate according to the above average browsing duration and the above total display duration;

[0202] When the above playback rate is greater than or equal to the first threshold and the above completion rate is greater than or equal to the second threshold, determine that the quality detection result of the above target candidate response data is passed.

[0203] In a possible implementation, when the above-mentioned processor 601 is used to determine multiple finely screened response data from the above-mentioned multiple roughly screened response data according to the content characterization information of each of the above-mentioned roughly screened response data, it is specifically used for:

[0204] Obtain the content characterization information of the target roughly screened response data; the target roughly screened response data is any one of the above-mentioned multiple roughly screened response data;

[0205] Detect the data content of the above-mentioned target roughly screened response data according to the content characterization information of the above-mentioned target roughly screened response data, and obtain a content detection result;

[0206] If the above content detection result is passed, determine the above target roughly screened response data as the finely screened response data.

[0207] In a possible implementation, the content characterization information of the above-mentioned target roughly screened response data includes: the display page of the content data of the above-mentioned target roughly screened response data, and the title information;

[0208] Wherein, when the above-mentioned processor 601 is used to detect the data content of the above-mentioned target rough screening response data according to the content characterization information of the above-mentioned target rough screening response data to obtain a content detection result, it is specifically used for:

[0209] Performing keyword extraction processing on the above-mentioned title information to obtain M first keywords, and performing keyword extraction processing on the above-mentioned display page to obtain N second keywords;

[0210] Determining K identical keywords among the above-mentioned M first keywords and the above-mentioned N second keywords;

[0211] When the ratio of K to P is greater than or equal to a third threshold, determining that the content detection result of the above-mentioned target rough screening response data is passed; M, N, and K are non-negative integers, and P is the maximum value of M and N.

[0212] In a possible implementation manner, the above-mentioned processor 601 is further used for:

[0213] Determining a text fragment set; the above-mentioned text fragment set includes multiple text fragments, and the above-mentioned target text fragment is any one of the above-mentioned multiple text fragments;

[0214] Determining multiple matching data pairs corresponding to each of the above-mentioned text fragments, and generating a training sample set according to the multiple matching data pairs corresponding to each of the above-mentioned text fragments;

[0215] Using the training samples in the above-mentioned training sample set to perform model training on an initial detection model to obtain a graphic-text consistency detection model.

[0216] In a possible implementation manner, when the above-mentioned processor 601 is used to determine the text fragment set, it is specifically used for:

[0217] Determining multiple historical search text fragments within a target time range according to the log data of the above-mentioned search engine;

[0218] Determining first historical search text fragments from the above-mentioned multiple historical search text fragments according to multiple search frequency intervals;

[0219] Determining second historical search text fragments from the above-mentioned multiple historical search text fragments according to multiple search categories;

[0220] Determining a text fragment set according to the above-mentioned first historical search text fragments and the above-mentioned second historical search text fragments.

[0221] In a specific implementation, the processor 601, storage device 602, and communication interface 603 described in the embodiments of the present application may execute the foregoing embodiments of the present application Figure 2 Or Figure 3The implementation manners described in the related embodiments of the provided data processing method can also be implemented in the embodiments of the present application. Figure 5 The implementation manners described in the related embodiments of the provided data processing apparatus will not be elaborated herein.

[0222] In several embodiments provided by the present application, it should be understood that the disclosed methods, apparatuses and systems can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there may be other division manners in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical or other forms.

[0223] In addition, it should be noted here that: The embodiments of the present application also provide a computer-readable storage medium, and a computer program executed by the aforementioned data processing apparatus is stored in the computer-readable storage medium, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the methods in the foregoing embodiments. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium related to the present application, please refer to the description of the method embodiments of the present application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located at one place, or alternatively, executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain system.

[0224] According to one aspect of the present application, there is provided a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can execute the methods in the foregoing embodiments. Therefore, it will not be elaborated here.

[0225] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0226] It can be understood that in the specific embodiments of the present application, relevant data such as response data, feedback data, and content data are involved. When the above embodiments of the present application are applied to specific products or technologies, the collection, use, and processing of relevant data need to comply with the relevant regulations and standards in the relevant regions.

[0227] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0228] It should be noted that the descriptions such as "first" and "second" involved in the embodiments of the present application are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" can explicitly or implicitly include at least one such feature.

[0229] The above-disclosed are only some embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.

Claims

1. A data processing method, characterized in that The method includes: Obtaining a plurality of candidate response data corresponding to the target text segment; the plurality of candidate response data is determined by the search engine in response to the search request for the target text segment, and the candidate response data includes content data and title information, and the content data includes one or more of images and videos; Determining a plurality of roughly screened response data from the plurality of candidate response data according to the quality feedback information of each candidate response data; each quality feedback information is determined according to the feedback data of the historical search object corresponding to the candidate response data for the data quality; Determining a plurality of finely screened response data from the plurality of roughly screened response data according to the content characterization information of each roughly screened response data; each content characterization information is determined according to the content data and title information of the corresponding roughly screened response data; Generating a plurality of matching data pairs according to the content data and title information of the plurality of finely screened response data; the content data and title information included in each matching data pair are matched.

2. The method according to claim 1, wherein The determining a plurality of roughly screened response data from the plurality of candidate response data according to the quality feedback information of each candidate response data includes: Obtaining the quality feedback information of the target candidate response data; the target candidate response data is any one of the plurality of candidate response data; Detecting the data quality of the target candidate response data according to the quality feedback information of the target candidate response data to obtain a quality detection result; If the quality detection result is passed, determining the target candidate response data as the roughly screened response data.

3. The method according to claim 2, wherein The quality feedback information of the target candidate response data includes: the click-through rate of the target candidate response data, the average browsing duration; Wherein, the detecting the data quality of the target candidate response data according to the quality feedback information of the target candidate response data to obtain a quality detection result includes: Obtaining the exposure times and the total display duration of the target candidate response data; Determining the playback rate according to the click-through rate and the exposure times, and determining the completion rate according to the average browsing duration and the total display duration; When the playback rate is greater than or equal to the first threshold and the completion rate is greater than or equal to the second threshold, determining that the quality detection result of the target candidate response data is passed.

4. The method according to claim 1, characterized in that, The determining a plurality of finely screened response data from the plurality of roughly screened response data according to the content characterization information of each roughly screened response data includes: Obtaining the content characterization information of the target roughly screened response data; the target roughly screened response data is any one of the plurality of roughly screened response data; Detecting the data content of the target roughly screened response data according to the content characterization information of the target roughly screened response data to obtain a content detection result; If the content detection result is passed, determining the target roughly screened response data as the finely screened response data.

5. The method according to claim 4, characterized in that The content characterization information of the target roughly screened response data includes: the display page of the content data of the target roughly screened response data, the title information; Among them, detecting the data content of the target rough screening response data according to the content characterization information of the target rough screening response data to obtain a content detection result includes: Performing keyword extraction processing on the title information to obtain M first keywords, and performing keyword extraction processing on the display page to obtain N second keywords; Determining K identical keywords among the M first keywords and the N second keywords; When the ratio of K to P is greater than or equal to a third threshold, determining that the content detection result of the target rough screening response data passes the detection; M, N, and K are non-negative integers, and P is the maximum value of M and N.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Determining a text fragment set; the text fragment set includes multiple text fragments, and the target text fragment is any one of the multiple text fragments; Determining multiple matching data pairs corresponding to each text fragment, and generating a training sample set according to the multiple matching data pairs corresponding to each text fragment; Using the training samples in the training sample set to train an initial detection model to obtain a graphic-text consistency detection model.

7. The method according to claim 6, wherein The determining the text fragment set includes: Determining multiple historical search text fragments within a target time range according to the log data of the search engine; Determining first historical search text fragments from the multiple historical search text fragments according to multiple search frequency intervals; Determining second historical search text fragments from the multiple historical search text fragments according to multiple search categories; Determining a text fragment set according to the first historical search text fragments and the second historical search text fragments.

8. A data processing device, characterized in that The device includes: An acquisition module, configured to acquire multiple candidate response data corresponding to a target text fragment; the multiple candidate response data are determined by the search engine in response to a search request for the target text fragment, and the candidate response data includes content data and title information, and the content data includes one or more of images and videos; A rough screening module, configured to determine multiple rough screening response data from the multiple candidate response data according to the quality feedback information of each candidate response data; each quality feedback information is determined according to the feedback data of the historical search object corresponding to the candidate response data for data quality; A fine screening module, configured to determine multiple fine screening response data from the multiple rough screening response data according to the content characterization information of each rough screening response data; each content characterization information is determined according to the content data and title information of the corresponding rough screening response data; A processing module, configured to generate multiple matching data pairs according to the content data and title information of the multiple fine screening response data; each matching data pair includes matching content data and title information.

9. A computer device, characterized in that, Includes: A processor, a storage device, and a communication interface, the processor, the communication interface, and the storage device are connected to each other, wherein the storage device stores executable program code, and the processor is configured to call the executable program code to implement the data processing method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, they are used to implement the data processing method described in any one of claims 1-7.

11. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they are used to implement the data processing method described in any one of claims 1-7.