Sample splicing method, device and equipment and computer readable storage medium

By gradually building a time window and storing data in external memory, the problems of real-time and resource consumption in the sample stitching scheme are solved, timely stitching and efficient output of samples are achieved, and the real-time and accuracy of the recommended model is improved.

CN120448800APending Publication Date: 2025-08-08TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410175583.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing sample stitching scheme cannot take into account the real-timeness of sample stitching and memory resource consumption, resulting in the recommendation model being unable to perceive object behavior in time, affecting the recommendation effect.

Method used

Sample splicing is performed using a progressive construction time window. By storing data to be spliced in external memory, memory consumption is reduced, and sample splicing is performed in the time window until the sample splicing is completed or the maximum number of time windows is reached.

Benefits of technology

It improves the real-time performance of sample construction and resource utilization efficiency, ensures that the recommendation model can perceive object behavior in a timely manner, and improves the accuracy of the recommendation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448800A_ABST
    Figure CN120448800A_ABST
Patent Text Reader

Abstract

The invention provides a sample splicing method, device and equipment and a computer readable storage medium, and the method comprises the steps: constructing a first time window under the condition of receiving display data; obtaining target data associated with the first request identifier based on the first time window; performing sample splicing based on the display data and the target data to obtain a spliced sample of the first time window; if the sample splicing result of the first time window is that the sample splicing is completed, stopping constructing the next time window, and generating a target sample identified by the first request according to the spliced sample of the first time window; or if the sample splicing result of the first time window is that the sample splicing is not completed, constructing a next time window, taking the next time window as a new first time window until the Nth time window, and generating a target sample identified by the first request according to the spliced sample of the Nth time window, the sample splicing result of the Nth time window is that the sample splicing is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more particularly, to a sample splicing method, apparatus, device, and computer-readable storage medium. Background Art

[0002] Recommendation systems achieve personalized and accurate recommendations by learning and simulating object behaviors. Therefore, they need to know the content recommended by the server and whether the object has a series of behaviors (such as exposure, clicks, plays, likes, favorites, and add-to-cart). The process of splicing the object behavior of the recommended content with the characteristics of the recommended items is called sample splicing.

[0003] In related technologies, there are two mainstream sample splicing solutions: offline sample splicing, which typically performs offline sample construction on a daily or hourly basis; and real-time sample construction, which primarily uses real-time processing frameworks to read data from message queues in real time and then store intermediate states in memory to complete the splicing. However, neither of these sample splicing solutions can balance the real-time nature of sample splicing and memory resource consumption. Summary of the Invention

[0004] The present application provides a sample splicing method, apparatus, device, and computer-readable storage medium, which can take into account both the real-time performance and memory resource consumption of sample splicing.

[0005] In a first aspect, a sample splicing method is provided, comprising:

[0006] Upon receiving display data, constructing a first time window, wherein the display data is associated with a first request identifier, the first request identifier corresponds to a first recommendation request, and the display data is display data of a resource recommended based on the first recommendation request;

[0007] acquiring target data associated with the first request identifier based on the first time window, the target data including at least one of reflow data and object feedback data, wherein the display data, reflow data, and object feedback data based on the same recommendation request are associated with the same request identifier;

[0008] Performing sample splicing based on the display data and target data associated with the first request identifier obtained based on the first time window to obtain a spliced sample corresponding to the first time window, and determining a sample splicing result corresponding to the first time window based on the spliced sample corresponding to the first time window;

[0009] If the sample splicing result corresponding to the first time window is that the sample splicing is completed, stop constructing the next time window, and generate the target sample corresponding to the first request identifier based on the splicing sample corresponding to the first time window; or if the sample splicing result corresponding to the first time window is that the sample splicing is not completed, construct the next time window, use the next time window as the new first time window, and execute: based on the first time window, obtain the target data associated with the first request identifier, until the Nth time window, and generate the target sample corresponding to the first request identifier based on the splicing sample corresponding to the Nth time window, wherein the sample splicing result corresponding to the Nth time window is that the sample splicing is completed, and N is less than N max , or, the Nth time window is the Nth max time windows, N max Indicates the maximum number of time windows allowed for builds.

[0010] In a second aspect, a sample splicing device is provided, comprising:

[0011] A construction unit 710 is configured to construct a first time window upon receiving display data, wherein the display data is associated with a first request identifier, the first request identifier corresponds to a first recommendation request, and the display data is display data of a resource recommended based on the first recommendation request;

[0012] an acquiring unit 720 configured to acquire target data associated with the first request identifier based on the first time window, the target data including at least one of reflow data and object feedback data, wherein the display data, reflow data, and object feedback data based on the same recommendation request are associated with the same request identifier;

[0013] a sample splicing unit 730 configured to perform sample splicing based on the display data and target data associated with the first request identifier obtained based on the first time window, obtain a spliced sample corresponding to the first time window, and determine a sample splicing result corresponding to the first time window based on the spliced sample corresponding to the first time window;

[0014] The sample generation unit 740 is configured to stop constructing the next time window if the sample splicing result corresponding to the first time window is that the sample splicing is completed, and generate the target sample corresponding to the first request identifier based on the spliced sample corresponding to the first time window; or if the sample splicing result corresponding to the first time window is that the sample splicing is not completed, construct the next time window, use the next time window as the new first time window, and execute: based on the first time window, obtain the target data associated with the first request identifier, until the Nth time window, and generate the target sample corresponding to the first request identifier based on the spliced sample corresponding to the Nth time window, wherein the sample splicing result corresponding to the Nth time window is that the sample splicing is completed, and N is less than N max , or, the Nth time window is the Nth max time windows, N max Indicates the maximum number of time windows allowed for builds.

[0015] In a third aspect, a computer device is provided, comprising a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the first aspect or its implementations.

[0016] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program, which enables a computer to execute the method in the above-mentioned first aspect or its various implementations.

[0017] In a fifth aspect, a computer program product is provided, comprising computer program instructions, wherein the computer program instructions enable a computer to execute the method in the above-mentioned first aspect or its various implementations.

[0018] In a sixth aspect, a computer program is provided, which, when executed on a computer, enables the computer to execute the method in the first aspect or its various implementations.

[0019] Based on the above technical solution, by progressively constructing a time window and then splicing samples based on the progressively constructed time window, the progressive output of samples is achieved, so that the vast majority of samples can be sent to the recommendation model for training in a timely manner, thereby improving the timeliness of real-time sample construction, enabling the recommendation model to perceive object behavior more timely, and making subsequent recommendation effects more accurate. In addition, the progressive splicing and output of samples does not require storing too much data in memory, which can reduce resource consumption. Therefore, the sample splicing solution of the embodiment of the present application can take into account both the timeliness of samples and memory resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1It is a schematic diagram of an application scenario applicable to an embodiment of the present application.

[0021] Figure 2 It is a schematic diagram of an offline sample construction scheme.

[0022] Figure 3 It is a schematic diagram of a real-time sample construction scheme.

[0023] Figure 4 This is a schematic flow chart of a sample splicing method provided in an embodiment of the present application.

[0024] Figure 5 This is a schematic diagram of backflow data processing provided in an embodiment of the present application.

[0025] Figure 6 This is a schematic diagram of object feedback data processing provided in an embodiment of the present application.

[0026] Figure 7 This is a schematic diagram of sample splicing processing provided in an embodiment of the present application.

[0027] Figure 8 This is a schematic diagram of sample splicing using a progressive time window provided in an embodiment of the present application.

[0028] Figure 9 This is a schematic block diagram of a sample splicing device provided in an embodiment of the present application.

[0029] Figure 10 This is a schematic block diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. With respect to the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0031] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.

[0032] In this application, unless otherwise specified, "at least one" means one or more, and "plurality" means two or more. Furthermore, "and / or" describes the association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects are in an "or" relationship.

[0033] It should be understood that the first, second, etc. descriptions appearing in the embodiments of the present application are only for illustration and distinction of the description objects, and there is no order. They do not represent any special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.

[0034] It should also be understood that the specific features, structures, or characteristics associated with the embodiments in the specification are included in at least one embodiment of the present application. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0035] In addition, the terms "include" and "have" and any variations thereof are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or elements is not necessarily limited to those steps or elements expressly listed, but may include other steps or elements not expressly listed or inherent to such process, method, product, or device.

[0036] The embodiments of the present application provide a sample splicing method, apparatus, device and computer-readable storage medium. Exemplarily, the sample splicing method can be executed by a computer device, wherein the computer device can be a terminal or a server. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, a smart speaker, a wearable smart device, a personal computer (PC), a smart car terminal and other devices. The terminal can also include a client, which can be a video client, a shopping application client, an entertainment and leisure application client or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms.

[0037] The embodiments of the present application can be applied to scenarios such as artificial intelligence, machine learning, and resource recommendation.

[0038] First, some nouns or terms that appear in the description of the embodiments of this application are explained as follows:

[0039] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0040] Computer Vision (CV): Computer vision is the science of enabling machines to "see." Specifically, it refers to machine vision techniques such as using cameras and computers to replace the human eye in identifying and measuring objects, and further processing the images to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0041] Speech Technology: Key technologies include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction. Large model technology is revolutionizing the development of speech technology. Pre-trained models such as WavLM and UniSpeech, which leverage the Transformer architecture, possess strong generalization and versatility, enabling them to effectively handle a wide range of speech processing tasks.

[0042] Natural language processing (NLP) is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; it also involves computer science and mathematics. The pre-training model, an important technology for model training in the field of artificial intelligence, is developed from the large language model (Large Language Model) in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0043] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by demonstration. Pretrained models are the latest development in deep learning, integrating these techniques.

[0044] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0045] Sample Engineering: Sample engineering is a crucial step in connecting online services and offline model training in recommendation systems. Its primary responsibilities include sample concatenation and business-related Extract, Transform, and Load (ETL) processing. In recommendation systems, sample engineering encompasses the entire process from accessing object features and object logs to model training, including sample processing, concatenation, and the generation of positive and negative samples.

[0046] Sample splicing: Recommendation systems achieve personalized and accurate recommendations by learning and simulating object behaviors. Therefore, they need to know the content recommended by the server and whether the object has a series of behaviors (such as exposure, clicks, plays, likes, favorites, and add-to-cart). The process of splicing the object behavior of the recommended content with the characteristics of the recommended items is generally called sample splicing.

[0047] Positive and negative samples: In the recommendation system, samples with only exposure data but no interaction data are negative samples. Conversely, samples with exposure and click interaction data are positive samples.

[0048] Feature crossover: Because model training occurs after online estimation, sample features may contain future information. This can cause inconsistencies between the sample features used during model training and those used during online estimation, impacting model performance.

[0049] Feature reflow: Through the real-time feature reporting process, the recommendation results' positive ranking data, the requested object profile, and contextual information are stored on disk along with the recommendation results through the reflow service. This ensures, to a certain extent, that the sample features used during model training are consistent with the feature data used during online estimation, preventing feature cross-talk.

[0050] Figure 1 A schematic diagram showing an application scenario of an embodiment of the present application.

[0051] like Figure 1 As shown, it includes a terminal 102 and a server 104. The terminal 102 communicates with the server 104 via a network. Optionally, the terminal 102 may include an APP front end, and the server 104 may include an APP backend server.

[0052] In some implementations, terminal 102 refers to a device that supports a variety of human-computer interaction methods, has internet access, typically runs various operating systems, and has strong processing capabilities. Terminal 102 may be, but is not limited to, a smartphone, tablet computer, portable laptop computer, desktop computer, wearable device, or vehicle-mounted device.

[0053] Optionally, in an embodiment of the present application, the terminal 102 may include a client, which may be a video client, a shopping application client, an entertainment and leisure application client, a social application client, a browser client, or an instant messaging client, etc.

[0054] In some implementations, server 104 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The server may also become a node of the blockchain. The server may be one or more servers. When there are multiple servers, there are at least two servers for providing different services, and / or there are at least two servers for providing the same service, such as providing the same service in a load balancing manner. This is not limited in the embodiments of the present application.

[0055] In some specific embodiments, the server 104 may be a server of a recommendation system.

[0056] Exemplarily, the network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), a 4G network, a 5G network, Bluetooth, Wi-Fi, or a call network.

[0057] like Figure 1 As shown, a data storage system 106 may also be included. The data storage system 106 can store data required by the server 104, such as return data, object feedback data, etc. The data storage system can be integrated on the server 104 or deployed on the cloud or other servers, which is not limited in this embodiment of the application.

[0058] In related technologies, there are two mainstream sample splicing solutions: offline construction solution and real-time construction solution.

[0059] Figure 2This is an offline sample construction solution based on Spark or Hive SQL. Spark and Hive SQL are big data processing and analysis tools, while HDFS is a distributed file system suitable for large-scale data processing. It can be used in conjunction with Spark and Hive SQL to provide efficient, reliable, and scalable data storage and processing services.

[0060] Specifically, after various data are saved on disk, offline sample construction is performed regularly with each day or hour as a time unit.

[0061] However, this offline sample construction solution has the following disadvantages:

[0062] 1. Accuracy is a potential risk. Because exposure data and playback data have a known order (i.e., exposure data must follow playback data), the corresponding playback and interaction behaviors at the end of each hour are likely to fall into the data table of the next hour. For example, if a video is sent to a subject at 10:58 and the subject clicks on the video at 10:59, but the video is played at 11:09, there will be a problem with the display data and the subject's feedback data being mismatched, affecting the accuracy of the sample.

[0063] 2. Poor real-time performance. Typically, hourly builds are the fastest. Furthermore, data storage delays typically range from 10 to 20 minutes. Furthermore, resolving the previous issue typically requires waiting for the next hour of playback and interactive data tables. Adding the computational time of offline sample building itself, the overall delay can exceed two hours.

[0064] Figure 3 It is a solution that uses the Flink stream processing framework to implement real-time sample construction.

[0065] like Figure 3 As shown, Flink receives multiple data streams simultaneously and stores their intermediate states in memory. Because reflow data, exposure data, playback data, and interactive data follow a clear sequence, a timer is set. Samples are then constructed based on the data obtained when the timer expires, and real-time samples are output.

[0066] However, the sample build solution using the Flink stream processing framework has disadvantages:

[0067] 1. Excessive resource usage and low task stability.

[0068] Specifically, due to the large amount of reflux data, Flink's storage of the intermediate state of the reflux data in memory will seriously consume memory and easily cause node failures. Therefore, additional resources are needed to solve this problem, which increases costs.

[0069] 2. Slow project iteration. All business logic is implemented within Flink tasks, which slows down iteration. Even simply adding a feature requires code modifications and logic adjustments, and restarting such a major task is costly.

[0070] In summary, traditional sample splicing solutions cannot take into account both the real-time performance and resource consumption of sample construction.

[0071] In order to address the shortcomings of traditional sample splicing solutions, the embodiments of the present application propose a sample splicing method that progressively constructs a time window for sample splicing, solving the problem that all logs need to wait for a fixed time before they can be spliced. This allows most samples to be sent to the recommendation model for training in a timely manner, improves the timeliness of real-time sample construction, enables the recommendation model to perceive object behavior more promptly, and makes subsequent recommendation effects more accurate.

[0072] Furthermore, external storage is introduced to store the data to be spliced. This shifts from performing sample splicing in memory to writing data to external storage independently in columns, simplifying the splicing logic for multiple data streams and reducing the use of memory state, thereby improving stability. Furthermore, feature lists can be used to control the feature data written. This way, when new feature data is needed, there is no need to restart the task or modify the code; only the feature list needs to be adjusted, reducing the cost of restarting the task. This also improves the efficiency of adding algorithm experiments and adding recommendation scenario models, reducing operation and maintenance costs and computing costs.

[0073] Furthermore, the sample splicing solution provided by the embodiments of this application can be used to optimize the real-time sample splicing process in sample engineering, further providing data support for model training of the recommendation system, and improving the timeliness and stability of sample data. As the quality of training samples improves, the recommendation effect of the recommendation model will also improve.

[0074] To facilitate understanding of the technical solutions of the embodiments of the present application, the technical solutions of the present application are described in detail below through specific embodiments. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the scope of protection of the embodiments of the present application. The embodiments of the present application include at least part of the following contents.

[0075] Each embodiment of the present application provides a sample splicing method, which can be executed by a terminal or a server, or by both the terminal and the server. The embodiments of the present application illustrate the sample splicing method by taking the server as an example.

[0076] See also Figures 4 to 8 ,in, Figure 4 is a schematic flow chart of a sample splicing method provided in an embodiment of the present application. Figure 5 This is a schematic diagram of backflow data processing provided by an embodiment of the present application. Figure 6 This is a schematic diagram of object feedback data processing provided by an embodiment of the present application. Figure 7 This is a sample splicing processing diagram provided in an embodiment of the present application. Figure 8 4 is a schematic diagram of a sample splicing method using a progressive time window provided in an embodiment of the present application. The method 400 may include at least some of the steps:

[0077] S401: Upon receiving display data (or exposure data), construct a first time window, wherein the display data is associated with a first request identifier, the first request identifier corresponds to a first recommendation request, and the display data is display data of a resource recommended based on the first recommendation request.

[0078] It should be understood that this application is not limited to specific recommendation scenarios, and may include but is not limited to: video recommendations, music recommendations, e-commerce recommendations, advertising recommendations, game recommendations, social media recommendations, etc.

[0079] In some embodiments, when a target user uses a certain media through a terminal and a recommendation node is triggered, a recommendation request is generated, which may include a request identifier. The terminal then sends the recommendation request to a recommendation device, which responds to the recommendation request and recommends the resource, thereby displaying or exposing the resource.

[0080] Optionally, the target object may be a related user who receives the recommended resource.

[0081] Optionally, the recommended resources may include but are not limited to videos, games, films, products, literary works, applications, etc.

[0082] In some embodiments, each recommendation request corresponds to a request identifier, and the display data, reflow data, and object feedback data (or feedback data) generated based on the same recommendation request are associated with the same request identifier. In this way, the associated data can be obtained based on the request identifier, and then sample splicing can be performed.

[0083] In some embodiments, the recommendation request may include object portrait information, context information, etc., so that the recommendation device can make accurate recommendations.

[0084] In some embodiments, the object feedback data may include playback data of the recommended resource or target object's interaction data with the recommended resource, etc. For example, the interaction data may include but is not limited to likes, favorites, forwarding, comments, and add-to-cart.

[0085] In some embodiments, the feedback type corresponding to the subject feedback data can be divided into positive feedback or negative feedback, where positive feedback indicates the target subject's active feedback or participation in the recommended resource, such as the target subject clicking, collecting, commenting on the recommended resource. Negative feedback indicates the target subject's negative feedback or dislike of the recommended resource, such as the target subject turning the page, jumping, unfollowing, or complaining about the recommended resource.

[0086] In some embodiments, the backflow data includes, for example, but is not limited to, at least one of the following: positive ranking data of the recommendation results, object portraits in the recommendation request, context information, and recommendation results.

[0087] In some embodiments, receiving presentation data may include:

[0088] The server receives presentation data from the terminal.

[0089] For example, after the server recommends a resource to the terminal based on the first recommendation request, it may generate display data of the resource. The terminal may send the display data to the server, where the display data is associated with the first request identifier.

[0090] In some embodiments, the method 400 further includes:

[0091] The server receives object feedback data associated with the first request identifier from the terminal.

[0092] For example, after a server recommends a resource to a terminal, the target user performs certain actions on the recommended resource, such as liking, adding to favorites, forwarding, or commenting on it, thereby generating object feedback data associated with the first recommendation request. The terminal can then send this object feedback data to the server. Alternatively, the target user may perform one or more feedback actions for the recommended resource, so a single request identifier can correspond to one or more pieces of object feedback data, such as play, like, or comment.

[0093] In some embodiments, the method 400 further includes:

[0094] The server writes the object feedback data and the reflow data associated with the first request identifier into the external memory.

[0095] That is, the embodiment of the present application can use an external memory to store object feedback data and reflow data.

[0096] Exemplarily, the external storage may include but is not limited to HBase.

[0097] Among them, the data in HBase is stored in the form of tables, each table consists of rows and columns. Therefore, all data associated with the same request identifier can be stored in a row, with the request identifier as the primary key (key). The object feedback data and return data associated with the request identifier are stored as values in different columns of the row. In this way, by querying different columns through the request identifier, the various data associated with the request identifier can be obtained, thereby naturally completing the splicing process. That is, the spliced data can be directly obtained from HBase according to the request identifier. Therefore, there is no need to occupy memory to store large amounts of data. When the data is needed, it can be read directly from external storage, which can reduce costs. In addition, the writing tasks of each data are independent of each other, and adding or modifying data does not affect each other.

[0098] For example, Figure 5 As shown, the request identifier is used as the rowKey, and the return data and object feedback data associated with the request identifier are placed in the same row.

[0099] Combine Figure 5 , the processing of the return flow data is explained. The processing process is implemented through a real-time stream computing framework (such as Flink). Taking the Flink stream computing framework as an example, it can include a connection unit (such as a Kafka conector source), a processing unit, and a writing unit (such as an HBase Sink). The connection unit is used for scalable and reliable streaming data transmission between the message queue and Flink, the processing unit is used to process the return flow data read from the message queue, and the writing unit is used to write the processed return flow data to an external storage device, such as HBase.

[0100] In some embodiments, due to the large amount of reflow data, for example, a single reflow data may be as large as 1MB, with a peak value of 2TB / hour. Directly writing the reflow data to external storage places a high pressure on input and output (I / O). In an embodiment of the present application, before writing the reflow data to external storage, the reflow data can be feature filtered by Flink's processing unit to filter out unnecessary reflow data. Optionally, the reflow data can also be compressed to further reduce storage pressure.

[0101] In some implementations, the reflow data can be filtered based on first configuration information, where the first configuration information is used to configure the reflow data to be filtered out or the reflow data to be retained. For example, the reflow configuration information can be a feature list that indicates the features of the retained reflow data. In this way, when new reflow data is needed, there is no need to restart the task or modify the code; only the feature list needs to be adjusted, which can reduce the cost of restarting the task.

[0102] Combine Figure 6 , the processing of object feedback data is explained. The processing process is implemented through a real-time stream computing framework (such as Flink). Taking the Flink stream computing framework as an example, it can include a connection unit (not shown), a processing unit and a writing unit (not shown), wherein the connection unit is used for scalable and reliable streaming data transmission between the message queue and Flink, the processing unit is used to process the object feedback data read from the message queue, and the writing unit is used to write the processed object feedback data to an external storage device, such as HBase.

[0103] In some embodiments, one request identifier may correspond to multiple pieces of object feedback data. When writing the object feedback data to the external memory, the object feedback data corresponding to one request identifier may be merged, for example, by using an append method to write the data to the external memory.

[0104] In some embodiments of the present application, the method 400 further includes:

[0105] Pre-partitioning the external memory;

[0106] Performing a hash operation on the first request identifier to obtain a hash result;

[0107] The object feedback data and reflow data associated with the first request identifier are written into the partition corresponding to the hash result.

[0108] For example, it is pre-divided into 1000 partitions, each partition corresponds to a partition identifier, such as 0 to 999, and then a hash operation is performed on the first request identifier (for example, 100) to obtain a hash result, and the data associated with the first request identifier is written into the partition corresponding to the hash result (for example, partition 100).

[0109] When storing data in an external memory, if the external memory is not pre-partitioned, it will initially have only one partition. New partitions will be created only when this partition is full. This can result in some partitions having too much data and others having too little data, leading to data skew. In an embodiment of the present application, by pre-partitioning the external memory and then performing a hash operation on the request identifier to determine which partition the data associated with the request identifier is written to, it is possible to ensure that data associated with different request identifiers falls relatively evenly within each partition, thereby preventing the data skew problem.

[0110] Since the display data and object feedback data have a known order, and the object feedback data comes after the display data, after receiving the display data associated with a request identifier, the server can trigger the start of the sample splicing process, such as building the first time window.

[0111] S402: Obtain target data associated with the first request identifier based on the first time window, where the target data includes at least one of reflow data and object feedback data, wherein the display data, reflow data, and object feedback data based on the same recommendation request are associated with the same request identifier.

[0112] S403: performing sample splicing based on the display data and target data associated with the first request identifier obtained based on the first time window to obtain a spliced sample corresponding to the first time window, and determining a sample splicing result corresponding to the first time window based on the spliced sample corresponding to the first time window;

[0113] S404: If the sample splicing result corresponding to the first time window is sample splicing completion, stop constructing the next time window, and generate a target sample corresponding to the first request identifier based on the spliced sample corresponding to the first time window; or

[0114] If the sample splicing result corresponding to the first time window is that the sample splicing is not completed, construct the next time window, use the next time window as the new first time window, and execute: based on the first time window, obtain the target data associated with the first request identifier, until the Nth time window, and generate the target sample corresponding to the first request identifier based on the splicing sample corresponding to the Nth time window, wherein the sample splicing result corresponding to the Nth time window is that the sample splicing is completed, and N is less than N max , or, the Nth time window is the Nth max time windows, N max Indicates the maximum number of time windows allowed for builds.

[0115] As mentioned above, the real-time sample construction process typically includes a fixed wait time for potential playback, interaction, and other feedback data, as playback and interaction behaviors always occur after exposure. To maximize the sample splicing rate, this time window is typically set relatively large, typically ten to twenty minutes, which impacts sample splicing efficiency and reduces the timeliness of real-time sample construction.

[0116] Through statistical analysis of the time distribution of splicing data, we found that although 99.99% of the behaviors occurred within ten minutes after exposure, 95% of the behaviors actually occurred within one minute after exposure. While waiting for ten minutes or more certainly increased the sample splicing rate, it also slowed down the output time of a large number of samples.

[0117] Therefore, in an embodiment of the present application, after receiving the display data, a first time window can be constructed, and then a sample splicing process can be performed based on the first time window. If the sample splicing is completed, the construction of the time window is stopped. If the sample splicing is not completed, the construction of the time window can continue, and then the sample splicing process is performed based on the next time window until the sample splicing is completed, or the maximum number of time windows allowed to be constructed is reached.

[0118] Therefore, in an embodiment of the present application, a time window can be constructed progressively, and then samples can be spliced based on the progressively constructed time window, thereby achieving progressive output of samples, so that most samples can be spliced in a timely manner and sent to the recommendation model for training in a timely manner, thereby improving the timeliness of real-time sample construction, enabling the recommendation model to perceive object behavior more timely, and making subsequent recommendation effects more accurate.

[0119] In some embodiments, the duration of the first time window can be set to the duration required for the sample splicing rate to reach a first preset threshold. max The total length of the time windows can be determined based on the time required for the splicing rate to reach the second preset threshold. The first preset threshold is less than the second preset threshold. Exemplarily, the first preset threshold is 95% and the second preset threshold is 99.99%. In this way, the splicing of most samples can be completed based on a shorter time window, thereby enabling timely output of samples. Samples that have not been spliced together only need to be spliced together based on the time window constructed later, without having to increase the length of the time window in order to improve the sample splicing rate, resulting in samples that can be spliced together in advance not being output in time, affecting the real-time performance of sample construction.

[0120] In some embodiments, the length of the time window created later is greater than the length of the time window created earlier, so that enough data can be collected to complete the splicing of samples.

[0121] In some embodiments, the first time window is implemented based on a timer, wherein the duration of the timer can be dynamically set or adjusted. Optionally, the start time of the next created time window is the same as the end time of the current time window, or in other words, the start time of the next timer is the end time of the current timer.

[0122] In some embodiments, acquiring target data associated with the first request identifier based on the first time window includes:

[0123] At the end of the first time window, the target data associated with the first request identifier is obtained, for example, the object feedback data and the reflow data associated with the first request identifier are obtained from an external memory.

[0124] Furthermore, sample splicing is performed based on the acquired target data associated with the first request identifier and the acquired display data associated with the first request identifier to determine a spliced sample corresponding to the first time window.

[0125] As mentioned above, when all data associated with a request identifier are stored in a row with the request identifier as the primary key, the data obtained from the external storage completes the splicing process. In this way, there is no need to occupy content to store a large amount of data, reducing the consumption of memory resources.

[0126] In some embodiments of the present application, determining, based on the spliced samples corresponding to the first time window, a sample splicing result corresponding to the time window (for example, whether sample splicing is completed or sample splicing is incomplete) may include:

[0127] According to whether the spliced samples corresponding to the first time window include object feedback data, it is determined whether the sample splicing result corresponding to the first time window is sample splicing completed or sample splicing incomplete.

[0128] For example, if the splicing sample corresponding to the first time window includes object feedback data, it is determined that the sample splicing result corresponding to the first time window is that the sample splicing is completed. The object feedback data here can be positive feedback data, such as object feedback data such as play, like, favorite, comment, etc., or it can also be reverse feedback data, such as feedback data such as page turning or jump. No matter which object feedback data is obtained, it can be considered that the feedback data of the object on the recommended resource has been obtained. Therefore, sample splicing can be performed based on the object feedback data and display data to obtain the target sample without constructing the next time window. Furthermore, the target sample can be used to construct a training sample and sent to the recommendation model for training in a timely manner, which improves the timeliness of real-time sample construction and enables the recommendation model to perceive object behavior more timely, making subsequent recommendation effects more accurate.

[0129] For another example, if the spliced samples corresponding to the first time window do not include object feedback data, it is determined that the sample splicing result corresponding to the first time window is incomplete. When the spliced samples corresponding to the first time window do not include object feedback data, it means that the object feedback data associated with the first request identifier has not been obtained based on the first time window. In this case, the sample splicing is considered incomplete and the object feedback data is missing. Therefore, the next time window can be constructed and sample splicing can be performed based on the target data obtained in the next time window.

[0130] In some embodiments, the method 400 further includes:

[0131] The type of the target sample corresponding to the first request identifier is determined according to whether the target sample corresponding to the first request identifier includes at least one of object feedback data and a feedback type corresponding to the object feedback data.

[0132] For example, if the target sample corresponding to the first request identifier does not include object feedback data, it is determined that the type of the target sample corresponding to the first request identifier is a negative sample. max If no object feedback data is obtained within a time window, it can be considered that the object has no feedback on the recommended resource. Therefore, the target sample can be used as a negative sample.

[0133] For another example, if the target sample corresponding to the first request identifier includes object feedback data, and the feedback type corresponding to the object feedback data is positive feedback, such as playing, liking, collecting, commenting, etc., it is determined that the type of the target sample corresponding to the first request identifier is a positive sample.

[0134] For another example, if the target sample corresponding to the first request identifier includes object feedback data, and the feedback type corresponding to the object feedback data is reverse feedback, such as page turning or jumping behavior, it is determined that the type of the target sample corresponding to the first request identifier is a negative sample.

[0135] In some embodiments of the present application, the method 400 further includes:

[0136] Whether to construct a next time window is determined according to whether the sample splicing result corresponding to the first time window is completed or incomplete.

[0137] For example, if the sample splicing result corresponding to the first time window is that the sample splicing is completed, then stop constructing the next time window. Alternatively, if the sample splicing result corresponding to the first time window is that the sample splicing is not completed, and the first time window is not the Nth max If the total duration of the X time windows that have been constructed does not reach the preset duration, the next time window can be constructed and sample splicing can be performed based on the next time window.

[0138] In some embodiments of the present application, the time window corresponding to the target sample is the Xth time window, wherein the Xth time window is the time window in which the first corresponding sample splicing result is the sample splicing completion result, or the Xth time window is the Nth time window. max time windows, or the total duration of the X time windows that have been constructed reaches the preset duration, that is, the waiting time for sample splicing reaches the preset duration.

[0139] In some embodiments, after generating the target sample corresponding to the first request identifier, the method 400 further includes:

[0140] A training sample is constructed according to the target sample corresponding to the first request identifier, and the training sample is input into the recommendation model for model training.

[0141] For example, if the target sample corresponding to the first request identifier does not include object feedback data, a negative sample can be constructed based on the target sample, and the negative sample can be input into the recommendation model for model training.

[0142] For another example, if the target sample corresponding to the first request identifier includes object feedback data, and the feedback type of the object feedback data is positive feedback, a positive sample can be constructed based on the target sample and input into the recommendation model for model training.

[0143] For another example, if the target sample corresponding to the first request identifier includes object feedback data, but the feedback type of the object feedback data is reverse feedback, a negative sample can be constructed based on the target sample and input into the recommendation model for model training.

[0144] Combine Figure 7 and Figure 8 , the sample splicing process based on time window is explained.

[0145] Specifically, after receiving the display data associated with the first request identifier, the splicing process can be started. For example, the sample state corresponding to the display data is initialized, for example, initialized to a negative sample, and a timer is registered. The timer corresponds to a time window. At the end of the time window, that is, when the timer ends, the object feedback data and the reflow data are obtained from the external memory based on the first request identifier, and then sample splicing is performed based on the obtained object feedback data, reflow data and display data. According to the sample splicing result, the sample state is updated. For example, when the sample splicing is completed and the feedback type corresponding to the object feedback data is positive feedback, the sample state is updated to a positive sample. Furthermore, the obtained real-time samples can be output to a message queue so that the recommendation model can obtain real-time samples from the message queue for training.

[0146] In an embodiment of the present application, samples may be spliced based on multi-level time windows and output.

[0147] For example, in Figure 8 In the example, after receiving the display data stream, a first time window (for example, 1 minute) is started. At the end of the first time window, object feedback data and return flow data are read from the external storage based on the request identifier of the display data stream, and then sample splicing is performed based on the obtained data:

[0148] Case 1: When sample splicing is completed (for the specific judgment method, refer to the relevant description of the above embodiment), the construction of the next time window is stopped.

[0149] Case 1-1: If the feedback type corresponding to the object feedback data in the spliced sample is positive feedback, the spliced sample is used to construct a positive sample, and the constructed positive sample is output to the message queue.

[0150] Case 1-2: If the feedback type corresponding to the object feedback data in the spliced sample is reverse feedback, the spliced sample is used to construct a negative sample, and the constructed negative sample is output to the message queue.

[0151] Case 2: Sample splicing is not completed, that is, the object feedback data is not obtained. In this case, the second time window (for example, 2 minutes) is started. At the end of the second time window, the object feedback data and reflow data are read from the external memory based on the request identifier of the display data stream, and then sample splicing is performed based on the obtained data. The process is similar and will not be repeated here.

[0152] When the number of created time windows reaches the maximum number of time windows allowed to be created, if the samples are still not spliced together, negative samples are constructed using the spliced samples and the constructed negative samples are output to the message queue.

[0153] In summary, the sample splicing solution of the embodiment of the present application solves the problem of having to wait a fixed amount of time for all logs to be spliced by introducing a progressive construction time window for sample splicing. This allows the vast majority of samples to be sent to the recommendation model for training in a timely manner, improving the timeliness of real-time sample construction, enabling the recommendation model to perceive object behavior more promptly, and making subsequent recommendations more accurate. Furthermore, the progressive splicing and output of samples eliminates the need to store too much data in memory, reducing resource consumption.

[0154] On the other hand, the introduction of external storage for the data to be spliced shifts from in-memory sample splicing to writing data to external storage independently in columns. This simplifies the splicing logic for multiple data streams and reduces the use of memory state, thereby improving stability. Furthermore, the written reflow data can be controlled through the feature list. This way, when new reflow data is needed, there is no need to restart the task or modify the code; only the feature list needs to be adjusted. This reduces the cost of task restarts, improves the efficiency of new algorithm experiments and the iteration of new recommendation scenario models, and reduces operation and maintenance costs and computing costs.

[0155] Therefore, the sample splicing solution provided by the embodiment of this application can optimize the real-time sample splicing process in the sample project, further provide data support for the model training of the recommendation system, and improve the timeliness and stability of the sample data. As the quality of the training samples improves, the recommendation effect of the recommendation model will also improve.

[0156] The specific embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, a variety of simple modifications can be made to the technical solution of the present application, and these simple modifications all fall within the scope of protection of the present application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.

[0157] It should also be understood that in the various method embodiments of the present application, the order of the sequence numbers of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. It should be understood that these sequence numbers can be interchanged where appropriate, so that the embodiments of the present application described can be implemented in an order other than those shown or described.

[0158] Combined with the following Figures 9 and 10 , describe in detail the device embodiments of the present application.

[0159] Figure 9 FIG is a schematic block diagram of a sample splicing device 700 according to an embodiment of the present application. Figure 9 As shown, the sample splicing device 700 includes:

[0160] A construction unit 710 is configured to construct a first time window upon receiving display data, wherein the display data is associated with a first request identifier, the first request identifier corresponds to a first recommendation request, and the display data is display data of a resource recommended based on the first recommendation request;

[0161] an acquiring unit 720 configured to acquire target data associated with the first request identifier based on the first time window, the target data including at least one of reflow data and object feedback data, wherein the display data, reflow data, and object feedback data based on the same recommendation request are associated with the same request identifier;

[0162] a sample splicing unit 730 configured to perform sample splicing based on the display data and target data associated with the first request identifier obtained based on the first time window, obtain a spliced sample corresponding to the first time window, and determine a sample splicing result corresponding to the first time window based on the spliced sample corresponding to the first time window;

[0163] The sample generation unit 740 is configured to stop constructing the next time window if the sample splicing result corresponding to the first time window is that the sample splicing is completed, and generate the target sample corresponding to the first request identifier based on the spliced sample corresponding to the first time window; or if the sample splicing result corresponding to the first time window is that the sample splicing is not completed, construct the next time window, use the next time window as the new first time window, and execute: based on the first time window, obtain the target data associated with the first request identifier, until the Nth time window, and generate the target sample corresponding to the first request identifier based on the spliced sample corresponding to the Nth time window, wherein the sample splicing result corresponding to the Nth time window is that the sample splicing is completed, and N is less than N max , or, the Nth time window is the Nth max time windows, N max Indicates the maximum number of time windows allowed for builds.

[0164] In some embodiments, the sample splicing device 700 further includes:

[0165] The first determining unit is configured to determine a type of the target sample corresponding to the first request identifier according to whether the target sample corresponding to the first request identifier includes at least one of object feedback data and a feedback type corresponding to the object feedback data.

[0166] In some embodiments, the first determining unit is further configured to:

[0167] If the target sample corresponding to the first request identifier does not include object feedback data, determining that the type of the target sample corresponding to the first request identifier is a negative sample; or

[0168] If the target sample corresponding to the first request identifier includes object feedback data, and the feedback type corresponding to the object feedback data is positive feedback, determining that the type of the target sample corresponding to the first request identifier is a positive sample; or

[0169] If the target sample corresponding to the first request identifier includes object feedback data, and the feedback type corresponding to the object feedback data is reverse feedback, it is determined that the type of the target sample corresponding to the first request identifier is a negative sample.

[0170] In some embodiments, the sample splicing device 700 further includes:

[0171] The second determining unit is configured to determine whether the sample splicing result corresponding to the first time window is completed or incomplete according to whether the spliced samples corresponding to the first time window include object feedback data.

[0172] In some embodiments, the second determining unit is further configured to:

[0173] If the spliced samples corresponding to the first time window include object feedback data, determining that the sample splicing result corresponding to the first time window is sample splicing completed; or

[0174] If the spliced samples corresponding to the first time window do not include object feedback data, it is determined that the sample splicing result corresponding to the first time window is incomplete.

[0175] In some embodiments, after generating the target sample corresponding to the first request identifier, the apparatus further includes:

[0176] A training unit is used to construct a training sample according to the target sample corresponding to the first request identifier, and input the training sample into the recommendation model for model training.

[0177] In some embodiments, the length of the later constructed first time window is greater than the length of the earlier constructed first time window.

[0178] In some embodiments, the acquiring unit 720 is further configured to:

[0179] When the first time window ends, the target data associated with the first request identifier is obtained from an external memory.

[0180] In some embodiments, in the external memory, the object feedback data and the reflow data associated with the first request identifier are stored in a row, the row having the first request identifier as a primary key, and obtaining the target data associated with the first request identifier from the external memory includes:

[0181] Obtain target data associated with the first request identifier from the row corresponding to the first request identifier.

[0182] In some embodiments, the sample splicing device 700 further includes:

[0183] A processing unit, configured to pre-partition the external memory;

[0184] Performing a hash operation on the first request identifier to obtain a hash result;

[0185] Write the data associated with the first request identifier into the partition corresponding to the hash result.

[0186] In some embodiments, the sample splicing device 700 further includes:

[0187] a receiving unit, configured to receive, from a terminal, reflow data associated with the first request identifier;

[0188] A processing unit is used to perform feature filtering on the reflow data associated with the first request identifier based on first configuration information, wherein the first configuration information is used to indicate features that need to be filtered out or retained in the reflow data; and write the reflow data associated with the first request identifier after feature filtering into the external memory.

[0189] It should be noted that the functions of each module in the sample splicing device 700 in the embodiment of the present application can correspond to the specific implementation methods of any embodiment in the above-mentioned method embodiments, and will not be repeated here.

[0190] Each unit in the sample splicing device 700 can be implemented in whole or in part by software, hardware, or a combination thereof. Each unit can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software so that the processor can call and execute the corresponding operations of each unit.

[0191] For example, the sample splicing device 700 can be integrated into a terminal or server that has a storage device and a processor and has computing power, or the sample splicing device 700 is the terminal or server. The terminal can be a smart phone, tablet computer, laptop computer, smart TV, smart speaker, wearable smart device, personal computer (PC) and other devices. The terminal can also include a client, which can be a video client, browser client or instant messaging client, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0192] Optionally, the present application also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0193] Figure 10A schematic structural diagram of a computer device provided in an embodiment of the present application, such as Figure 10 As shown, computer device 800 may include: a communication interface 801, a memory 802, a processor 803, and a communication bus 804. Communication interface 801, memory 802, and processor 803 communicate with each other via communication bus 804. Communication interface 801 is used for data communication between device 800 and external devices. Memory 802 may be used to store software programs and modules. Processor 803 executes software programs and modules stored in memory 802, such as the software programs for corresponding operations in the aforementioned method embodiments.

[0194] Optionally, the processor 803 may call software programs and modules stored in the memory 802 to perform the following operations:

[0195] Upon receiving display data, constructing a first time window, wherein the display data is associated with a first request identifier, the first request identifier corresponds to a first recommendation request, and the display data is display data of a resource recommended based on the first recommendation request;

[0196] acquiring target data associated with the first request identifier based on the first time window, the target data including at least one of reflow data and object feedback data, wherein the display data, reflow data, and object feedback data based on the same recommendation request are associated with the same request identifier;

[0197] Performing sample splicing based on the display data and target data associated with the first request identifier obtained based on the first time window to obtain a spliced sample corresponding to the first time window, and determining a sample splicing result corresponding to the first time window based on the spliced sample corresponding to the first time window;

[0198] If the sample splicing result corresponding to the first time window is that the sample splicing is completed, stop constructing the next time window, and generate the target sample corresponding to the first request identifier based on the splicing sample corresponding to the first time window; or if the sample splicing result corresponding to the first time window is that the sample splicing is not completed, construct the next time window, use the next time window as the new first time window, and execute: based on the first time window, obtain the target data associated with the first request identifier, until the Nth time window, and generate the target sample corresponding to the first request identifier based on the splicing sample corresponding to the Nth time window, wherein the sample splicing result corresponding to the Nth time window is that the sample splicing is completed, and N is less than N max , or, the Nth time window is the Nth max time windows, N max Indicates the maximum number of time windows allowed for builds.

[0199] Optionally, the computer device 800 may be integrated into a terminal or server that has storage and a processor and has computing capabilities, or the computer device 800 may be the terminal or server. The terminal may be a smartphone, tablet computer, laptop computer, smart TV, smart speaker, wearable smart device, personal computer, or other device. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0200] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding processes in the above-mentioned methods in the embodiments of this application. For the sake of brevity, it is not further described here.

[0201] This application also provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes of the above-described methods in the embodiments of this application. For the sake of brevity, these processes are not further described here.

[0202] This application also provides a computer program comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes of the above-described methods in the embodiments of this application. For the sake of brevity, these processes are not further described here.

[0203] It should be understood that the processor of the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by hardware integrated logic circuits in the processor or software instructions. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly implemented as a hardware decoding processor, or can be implemented by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0204] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0205] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0206] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0207] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0208] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0209] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0210] In addition, each functional unit in the embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0211] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.

[0212] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A sample splicing method, characterized in that: include: Upon receiving display data, constructing a first time window, wherein the display data is associated with a first request identifier, the first request identifier corresponds to a first recommendation request, and the display data is display data of a resource recommended based on the first recommendation request; acquiring target data associated with the first request identifier based on the first time window, the target data including at least one of reflow data and object feedback data, wherein the display data, reflow data, and object feedback data based on the same recommendation request are associated with the same request identifier; Performing sample splicing based on the display data and target data associated with the first request identifier obtained based on the first time window to obtain a spliced sample corresponding to the first time window, and determining a sample splicing result corresponding to the first time window based on the spliced sample corresponding to the first time window; If the sample splicing result corresponding to the first time window is that the sample splicing is completed, stop constructing the next time window, and generate the target sample corresponding to the first request identifier based on the splicing sample corresponding to the first time window; or if the sample splicing result corresponding to the first time window is that the sample splicing is not completed, construct the next time window, use the next time window as the new first time window, and execute: based on the first time window, obtain the target data associated with the first request identifier, until the Nth time window, and generate the target sample corresponding to the first request identifier based on the splicing sample corresponding to the Nth time window, wherein the sample splicing result corresponding to the Nth time window is that the sample splicing is completed, and N is less than N max , or, the Nth time window is the Nth max time windows, N max Indicates the maximum number of time windows allowed for builds.

2. The method according to claim 1, characterized in that The method further comprises: The type of the target sample corresponding to the first request identifier is determined according to whether the target sample corresponding to the first request identifier includes at least one of object feedback data and a feedback type corresponding to the object feedback data.

3. The method according to claim 2, characterized in that The determining, according to whether the target sample corresponding to the first request identifier includes at least one of object feedback data and a feedback type corresponding to the object feedback data, the type of the target sample corresponding to the first request identifier includes: If the target sample corresponding to the first request identifier does not include object feedback data, determining that the type of the target sample corresponding to the first request identifier is a negative sample; or If the target sample corresponding to the first request identifier includes object feedback data, and the feedback type corresponding to the object feedback data is positive feedback, determining that the type of the target sample corresponding to the first request identifier is a positive sample; or If the target sample corresponding to the first request identifier includes object feedback data, and the feedback type corresponding to the object feedback data is reverse feedback, it is determined that the type of the target sample corresponding to the first request identifier is a negative sample.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: According to whether the spliced samples corresponding to the first time window include object feedback data, it is determined whether the sample splicing result corresponding to the first time window is sample splicing completed or sample splicing incomplete.

5. The method according to claim 4, characterized in that Determining, according to whether the spliced samples corresponding to the first time window include object feedback data, whether the sample splicing result corresponding to the first time window is completed or incomplete, includes: If the spliced samples corresponding to the first time window include object feedback data, determining that the sample splicing result corresponding to the first time window is sample splicing completed; or If the spliced samples corresponding to the first time window do not include object feedback data, it is determined that the sample splicing result corresponding to the first time window is incomplete.

6. The method according to any one of claims 1 to 3, characterized in that After generating the target sample corresponding to the first request identifier, the method further includes: A training sample is constructed according to the target sample corresponding to the first request identifier, and the training sample is input into the recommendation model for model training.

7. The method according to any one of claims 1 to 3, characterized in that The length of the later constructed first time window is greater than the length of the earlier constructed first time window.

8. The method according to any one of claims 1 to 3, characterized in that The acquiring, based on the first time window, target data associated with the first request identifier includes: When the first time window ends, the target data associated with the first request identifier is obtained from an external memory.

9. The method according to claim 8, characterized in that In the external memory, the object feedback data and the reflow data associated with the first request identifier are stored in a row, the row having the first request identifier as a primary key, and obtaining the target data associated with the first request identifier from the external memory includes: Obtain target data associated with the first request identifier from the row corresponding to the first request identifier.

10. The method according to claim 8, characterized in that The method further comprises: Pre-partitioning the external memory; Performing a hash operation on the first request identifier to obtain a hash result; Write the data associated with the first request identifier into the partition corresponding to the hash result.

11. The method according to claim 8, characterized in that The method further comprises: receiving, from the terminal, the return flow data associated with the first request identifier; Performing feature filtering on the reflow data associated with the first request identifier based on first configuration information, where the first configuration information is used to indicate features to be filtered out or retained in the reflow data; The reflow data associated with the first request identifier after feature filtering is written into the external memory.

12. A sample splicing device, characterized in that: include: a constructing unit, configured to construct a first time window upon receiving display data, wherein the display data is associated with a first request identifier, the first request identifier corresponds to a first recommendation request, and the display data is display data of a resource recommended based on the first recommendation request; an acquiring unit, configured to acquire target data associated with the first request identifier based on the first time window, the target data comprising at least one of reflow data and object feedback data, wherein the display data, reflow data, and object feedback data based on the same recommendation request are associated with the same request identifier; a sample splicing unit, configured to perform sample splicing based on the display data and target data associated with the first request identifier obtained based on the first time window, obtain a spliced sample corresponding to the first time window, and determine a sample splicing result corresponding to the first time window based on the spliced sample corresponding to the first time window; The sample generation unit is configured to stop constructing the next time window if the sample splicing result corresponding to the first time window is that the sample splicing is completed, and generate the target sample corresponding to the first request identifier based on the spliced sample corresponding to the first time window; or, if the sample splicing result corresponding to the first time window is that the sample splicing is not completed, construct the next time window, use the next time window as the new first time window, and execute: based on the first time window, obtain the target data associated with the first request identifier, until the Nth time window, and generate the target sample corresponding to the first request identifier based on the spliced sample corresponding to the Nth time window, wherein the sample splicing result corresponding to the Nth time window is that the sample splicing is completed, and N is less than N max , or, the Nth time window is the Nth max time windows, N max Indicates the maximum number of time windows allowed for builds.

13. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to execute the sample splicing method according to any one of claims 1 to 11 by calling the computer program stored in the memory.

14. A computer-readable storage medium, characterized in that Used to store a computer program, wherein the computer program causes a computer to execute the method according to any one of claims 1 to 11.