Data processing method and system and electronic equipment
By allocating memory shards to users and establishing an address mapping table, the timeliness and operational complexity issues in real-time feature data processing are resolved, achieving efficient data processing with in-memory computing integration and improving the read/write capability and response speed of feature data.
Patent Information
- Application Number
- CN202511125415.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies face limitations in real-time feature data processing in internet search, advertising, and recommendation scenarios, including timeliness limitations and high operational complexity. This is primarily due to redundant computation and long-link transmission delays caused by the separation of the streaming computing engine from the storage layer.
By allocating memory slices to users and marking memory address information on the memory slices, a strong association is established between user marking information and memory addresses, enabling memory-level in-memory computing and data processing. The address mapping table is used to quickly locate memory slices for data writing and reading, avoiding redundant storage and long-link transmission.
It significantly improves the ability to read and write real-time feature data, reduces the consumption of computing and storage resources, achieves millisecond-level response capability, and improves the processing efficiency and accuracy of feature data.
Smart Images

Figure CN121029082A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and more specifically, to a data processing method, system and electronic device. Background Art
[0002] In the scenarios of Internet search, advertising, and recommendation services, user real-time feature data refers to a set of dynamic attributes that can reflect the current status and behavior of users. User real-time feature data needs to be collected, processed, and supplied with millisecond-level timeliness, and directly serves model inference and policy decision-making in scenarios such as search results, advertising placement, and content recommendation. Its core value lies in capturing key information such as users' instantaneous interest preferences, device environment characteristics, and interaction trajectories, providing a highly fresh input basis for recommendation algorithms, and thus significantly improving the accuracy and conversion efficiency of recommendations. Since the traditional offline data processing mode can no longer meet the response speed requirements of modern intelligent services, after a user completes a search or click action, it is necessary to complete the complete closed-loop of user behavior feature extraction, prediction, and policy matching recommendation within an extremely short time. Delays in any link may lead to a decline in user experience or the loss of business opportunities. Therefore, building an efficient real-time data processing link has become the key infrastructure to support business competitiveness.
[0003] The current mainstream technical architecture in the industry mainly relies on a streaming computing engine, cooperates with a behavior feature data collection tool to implement message queue transmission, and uses a general storage medium to save the calculation results. Although this combination can basically meet the business requirements, there are still many structural defects: Feature calculation based on a sliding window will trigger a large number of redundant feature recalculations; the physical separation of the calculation layer and the storage layer causes data to experience multiple disk write operations, increasing the processing delay. These technical pain points together lead to the core problems of the traditional real-time feature data processing link, such as the timeliness ceiling and high operation and maintenance complexity. Summary of the Invention
[0004] The present invention provides a data processing method, system and electronic device for efficiently processing user feature data in real time, improving the timeliness of processing user feature data and reducing operation and maintenance costs.
[0005] According to the first aspect of the present application, a data processing method is provided, and the method includes: Pre-allocate memory shards for a user, and the memory shards mark memory address information; Collect behavior data carrying user marking information; Extract the feature data of the behavior data, and the feature data carries the user marking information; Based on the user marking information of the feature data and the memory address information of the memory shards, write the feature data into the corresponding memory shards; Obtain guidance information for the data to be read, wherein the guidance information includes user tagging information; Based on the user tag information of the guidance information and the memory address information of the memory slice, the data to be read is obtained from the corresponding memory slice.
[0006] Understandably, by pre-allocating independent memory slices with clearly defined memory addresses to users, a data access mechanism based on a strong correlation between user tagging information and memory address information is constructed. During the data writing phase, user tagging information carried by the feature data extracted from behavioral data is accurately written to the corresponding memory slice, completely eliminating the time-consuming losses caused by redundant intermediate storage and long-link transmissions resulting from reliance on different tools in traditional architectures. During the data reading phase, guided by information containing user tagging information, the target memory slice can be quickly located to complete data reading, avoiding resource waste caused by full scans or complex index calculations. The memory-level integrated storage and computing design in this solution achieves efficient writing and low-latency reading of feature data, significantly improving the read / write capability of real-time feature data while reducing the consumption of computing and storage resources.
[0007] Optionally, the method further includes: Establish an address mapping table for user and memory address information for the memory slices; The process of writing the feature data into the corresponding memory slice based on the user tagging information of the feature data and the memory address information of the memory slice includes: Construct a write request for the feature data based on the feature data; The address mapping table is queried according to the user tag information of the feature data to obtain the memory address information that matches the user tag information of the feature data, and the memory slice to be written is determined according to the memory address information that matches the user tag information of the feature data. The write request is forwarded to the memory shard to be written, and the write operation of the feature data is completed in the memory shard to be written.
[0008] Optionally, an address mapping table for user and memory address information is established for the memory slices; The process of retrieving the data to be read from the corresponding memory segment based on the user tagging information of the guidance information and the memory address information of the memory segment includes: Construct a read request for the data to be read based on the guidance information; The address mapping table is queried according to the user tag information of the guidance information to obtain the memory address information that matches the user tag information of the guidance information, and the memory slice to be read is determined according to the memory address information that matches the user tag information of the guidance information. The read request is forwarded to the memory shard to be read, and the read operation of the data to be read is completed in the memory shard to be read.
[0009] Understandably, by constructing an address mapping table between user and memory address information, efficient and accurate access to feature data is achieved. During the write phase, the target address is determined based on user tagging information. The address mapping table is traversed to quickly match the memory slice to be written to, and the write request is forwarded to that slice to complete the data storage. This completely avoids the redundant computation and timeliness bottlenecks caused by cross-tool transmission and intermediate storage in traditional architectures. During the read phase, the read address is also obtained based on user tagging information. The address mapping table is used to quickly locate the memory slice to be read and execute the read operation, eliminating the need for full scans or complex index calculations. This address-mapping-based targeted access mechanism significantly shortens the data flow path, eliminating resource waste caused by the separation of storage and computation, and significantly reducing I / O input / output load and computational overhead. This forms a highly efficient memory-level integrated computing architecture, greatly improving the real-time processing capability and response speed of feature data.
[0010] Optionally, the extraction of feature data from the behavioral data includes: Different user types are preset, and corresponding feature indicators are set for each user type; Based on the user tagging information carried by the behavioral data, the user type corresponding to the behavioral data is determined; Determine the corresponding feature indicators based on the identified user type; Based on the determined feature indicators, extract the feature data corresponding to the behavioral data.
[0011] Understandably, by pre-setting different user types and their associated feature indicators, refined feature extraction of behavioral data is achieved. Based on the user tagging information carried by the behavioral data, the user type is identified, and the corresponding feature indicator set is matched for targeted extraction. This shifts the feature data extraction from general full extraction to personalized adaptation extraction, avoiding the blind processing of full feature data and ensuring the accurate capture of key feature data. This significantly improves the effectiveness and relevance of feature data, provides high-quality structured feature data for subsequent analysis, and reduces ineffective computational overhead.
[0012] Optionally, the method further includes: Several fragment copies are pre-configured for the memory fragments; After writing the feature data into the corresponding memory shard, the written feature data is also synchronized to the shard copy corresponding to the memory shard. Based on the user tagging information of the guidance information and the memory address information of the memory slice, the data to be read is obtained from the corresponding memory slice, including: Based on the user tagging information of the guidance information and the memory address information of the memory shard, the data to be read is obtained from the corresponding memory shard and / or the corresponding shard copy.
[0013] Understandably, by pre-configuring several shard replicas for the memory shards and achieving data synchronization, the stability and fault tolerance of feature data processing are significantly enhanced. After feature data is written to the memory shard, the written feature data is automatically synchronized to the corresponding shard replica, forming a master-slave consistent redundant data storage structure. When performing data read operations, data can be flexibly selected to be obtained from the memory shard or any shard replica according to actual needs. This not only effectively avoids the risk of single point of failure and ensures the persistent storage and continuous reading and writing of feature data, but also optimizes resource utilization through read-write separation. Read requests can be distributed to multiple shard replicas, which not only alleviates the pressure on the memory shards, but also improves the efficiency of concurrent reads, while maintaining consistent data protection, providing highly reliable and low-latency technical support for real-time feature data processing.
[0014] Optionally, the memory shard comprises several blocks, and the several blocks are distributed to store the feature data written in the memory shard, with each block synchronously writing the occurrence time of the stored feature data.
[0015] Understandably, by dividing memory into multiple distributed storage blocks and synchronously recording the occurrence time of feature data, efficient management and accurate traceability of feature data are achieved. Each block adopts a distributed storage architecture, which can handle the writing and storage of large-scale feature data, significantly improving data processing throughput. At the same time, synchronously writing the occurrence time of feature data ensures that feature data in different blocks have a unified time base, which not only guarantees the temporal consistency of feature data, but also provides accurate time stamps for subsequent time-based analysis, effectively optimizing the utilization efficiency of memory resources, reducing the pressure on single-point storage, and enhancing data reliability and traceability, providing efficient technical support for real-time feature data calculation and historical backtracking analysis.
[0016] Optionally, writing the feature data into the corresponding memory slice includes: According to the occurrence time of the feature data, the feature data and the occurrence time of the feature data are written into the corresponding memory slice block; If the feature data of the current occurrence is written to the tail block of the memory slice, then the feature data of the next occurrence is written to the head block of the memory slice, overwriting the data already stored in the head block.
[0017] Understandably, during the writing process, the occurrence time of feature data is synchronously recorded and stored in blocks, forming a time-series data organization. When the data writing reaches the end of the block, the data from the next occurrence time will automatically wrap around to the beginning of the block for overwriting. This circular buffer design avoids the risk of unlimited memory growth and ensures that the latest timely feature data is always retained. Through refined time dimension management and dynamic overwriting strategies, the time lag problem caused by long link transmission in traditional solutions is effectively solved. At the same time, the compact memory layout reduces storage redundancy, enabling feature data to complete the entire process from collection to storage with minimal latency, providing millisecond-level response capability for real-time recommendation decisions.
[0018] Optionally, the guidance information for the data to be read may also include feature indicators and time range; The step of obtaining the data to be read from the corresponding memory slice includes: Based on the time range, determine the block to be read in the corresponding memory slice; Based on the feature indicators of the guidance information, the feature data in the block to be read is calculated to obtain the data to be read.
[0019] Understandably, the guidance information includes feature indicators and time ranges, constructing a feature data reading and calculation system that combines accuracy and flexibility. After quickly locating the memory slice to be read based on user tagging information, the data range to be processed can be accurately defined according to the preset time range. Subsequently, the feature indicators are combined to perform targeted calculations on the feature data within the time range, realizing the efficient transformation from raw feature data to the required feature indicators. This avoids the invalid traversal of the entire feature data and flexibly adapts to the analysis needs of different time granularities through the sliding window mechanism, significantly improving the reading efficiency and computational performance of feature data, and providing fine-grained and timely data support capabilities for real-time recommendation decisions.
[0020] According to a second aspect of this application, a data processing system is provided, the system comprising: The memory shard allocation module is used to pre-allocate memory shards for users, and the memory shards are marked with memory address information; The data collection module is used to collect behavioral data carrying user tagging information; The extraction module is used to extract feature data from the behavioral data, wherein the feature data carries the user tagging information; The writing module is used to write the feature data into the corresponding memory segment based on the user tag information of the feature data and the memory address information of the memory segment; The guidance information acquisition module is used to acquire guidance information for the data to be read, wherein the guidance information includes user tag information; The reading module is used to obtain the data to be read from the corresponding memory segment based on the user tag information of the guidance information and the memory address information of the memory segment.
[0021] According to a third aspect of this application, an electronic device is provided, comprising: Memory, used to store one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements the data processing method described in the first aspect above.
[0022] According to a fourth aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data processing method described in the first aspect above.
[0023] Based on any of the above aspects, embodiments of this application provide a data processing method, system, electronic device, and storage medium that pre-allocates memory shards for a user, the memory shards being labeled with memory address information; collects behavioral data carrying user labeling information; extracts feature data from the behavioral data, the feature data carrying the user labeling information; writes the feature data into the corresponding memory shard based on the user labeling information of the feature data and the memory address information of the memory shard; obtains guidance information for data to be read, wherein the guidance information includes user labeling information; and obtains the data to be read from the corresponding memory shard based on the user labeling information of the guidance information and the memory address information of the memory shard. This application can achieve the following benefits: • Reduce redundant storage and overcome timeliness bottlenecks: By pre-allocating memory shards for users and establishing a strong correlation between user tagging information and memory address information, the traditional process of feature data processing is completely restructured. Compared to the long-link storage architecture of existing technologies that relies on multiple tools and goes through multiple levels of transfer such as message queues, temporary storage, batch processing storage, and final storage, this solution directly directs the writing of feature data to the target memory shard, avoiding redundant intermediate storage links and eliminating duplicate memory overhead caused by cross-tool data transfer; at the same time, since feature data does not need to experience the read / write latency of external storage systems, the time from data collection to writing is greatly shortened, significantly improving the timeliness of real-time feature data read and write.
[0024] • Precise memory shard location for optimized computational performance: When user data needs to be retrieved, the system can directly map to the corresponding memory shard based on the user tag information in the guidance message, eliminating the need to traverse the distributed storage cluster. This fundamentally solves the long-chain problem caused by full scanning of the storage medium or complex index queries in traditional solutions. Thanks to the physical isolation characteristics of memory shards, a single read operation only involves local feature data within the target memory shard, significantly reducing I / O input / output load and CPU processor computing resource consumption, achieving efficient and lightweight data access.
[0025] • Constructing an integrated storage and computing system to achieve computing and storage fusion: Unlike the frequent data migration caused by the separation of storage and computing in existing technologies, this application enables the writing, storage, and computing of feature data to all occur within a pre-defined memory shard, forming a closed loop of integrated storage and computing. This allows operations such as writing, computing, transforming, and reading feature data to directly affect the original data in the memory shard, avoiding the overhead of data copying across storage media, ensuring the accuracy of feature data during computing, and enabling the reading and writing of complex feature data with lower latency, providing millisecond-level response capabilities for real-time user recommendations and other decisions. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is an illustrative application scenario diagram of the data processing method provided in this embodiment.
[0028] Figure 2 This is a flowchart of a data processing method provided in this embodiment.
[0029] Figure 3 This is a flowchart for extracting feature data provided in this embodiment.
[0030] Figure 4 This is a flowchart for obtaining a memory slice to be written, as provided in this embodiment.
[0031] Figure 5 This is a flowchart of a write operation provided in this embodiment.
[0032] Figure 6 This embodiment provides a flowchart for obtaining a memory slice to be read.
[0033] Figure 7This is a flowchart of a read operation provided in this embodiment.
[0034] Figure 8 This is a schematic diagram of the functional modules of a data processing system provided in this embodiment.
[0035] Figure 9 This is a schematic diagram of the structure of the electronic device provided in this embodiment. Detailed Implementation
[0036] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this application. To better illustrate the following embodiments, some components in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product; it is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application.
[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0039] In search, advertising, and recommendation scenarios, user behavior prediction and recommendation strategies heavily rely on the efficient computation and storage of real-time user feature data. This places stringent demands on the timeliness and computational power of feature data processing. Currently, the industry generally adopts a traditional architecture based on streaming computing engines, feature data acquisition tools, and related script statistics. This architecture disperses feature data across traditional storage media and performs sliding window computation through streaming computing engines. However, this approach has significant drawbacks: firstly, the long read / write link architecture leads to cumulative data transmission delays, and timeliness is limited by the processing cycle of each tool; secondly, sliding window computation causes a large amount of repetitive computation; and the separation of storage and computation in the traditional architecture generates redundant intermediate storage, further exacerbating resource waste and performance bottlenecks.
[0040] This embodiment provides a technical solution that can solve the above problems. The specific implementation of this application will be described in detail below with reference to the accompanying drawings.
[0041] An exemplary diagram illustrating an application scenario of a data processing method provided in this application embodiment is shown below. Figure 1 As shown, the application scenario includes at least a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has data writing, data calculation, and data reading functions; the terminal 200 has data extraction, data writing, data calculation, and data reading functions.
[0042] Understandably, the server 100 can be an independent electronic device or a cluster of multiple electronic devices; the terminal 200 can be a smartphone terminal, personal computer, tablet computer, vehicle terminal, etc., but is not limited to these.
[0043] In one possible implementation, server 100 and terminal 200 may each execute a data processing method provided in the embodiments of this application. Alternatively, a data processing method provided in the embodiments of this application may be executed partly in server 100 and partly in terminal 200.
[0044] like Figure 2 As shown, this embodiment provides a data processing method, which can be further divided into the following steps: S100: Pre-allocate memory slices for the user, wherein the memory slices are marked with memory address information; In this embodiment, the user can be any type of user using the platform. When users use the platform, they can generate different behavioral data. The feature data corresponding to this behavioral data can guide the search, advertising, and recommendation services in the platform, enabling the platform to obtain the user's viewing and / or operation preferences based on the feature data, thereby generating relevant search terms, related advertising services, and related recommended content, which can improve the user stickiness of the platform.
[0045] In this embodiment, a memory slice is allocated to each user in the platform, that is, each user has a corresponding memory slice. Subsequently, the user's feature data and occurrence time can be stored in the memory slice, so that the feature data and occurrence time of each user are stored in an orderly manner, improving the efficiency of writing and reading.
[0046] Understandably, when the corresponding user feature data is obtained, it is necessary to find the matching memory slice in order to complete the writing operation. In this embodiment, by marking the memory address information in each memory slice as the location identifier of each memory slice, the corresponding memory slice can be found according to the memory address information in subsequent write or read operations, and the feature data of the corresponding memory slice can be written or read.
[0047] Preferably, in order to accommodate a large number of users on the platform, this embodiment will set up several service containers based on the concept of distributed storage. Each service container contains a large number of memory shards. When the number of users changes, the access service containers and / or the number of memory shards used in the service containers can be flexibly adjusted, thereby quickly and flexibly responding to changing scenarios and improving the reliability of user feature data processing.
[0048] Specifically, in this embodiment, it further includes: establishing an address mapping table for user and memory address information for the memory slice; In this embodiment, to avoid directly traversing a large number of memory slices and their corresponding memory address information when searching for the corresponding memory slice, it is necessary to establish an address mapping table between users and memory address information based on the memory address information of the memory slices. This address mapping table can record user-related information and corresponding memory address information, so that when searching for the corresponding memory slice, only the address mapping table needs to be traversed to quickly find the corresponding memory address information, thereby finding the corresponding memory slice and improving the efficiency of the query.
[0049] Preferably, this embodiment further includes a main service component, in which the address mapping table is established and stored.
[0050] S200: Collect behavioral data carrying user tagging information; In this embodiment, user behavior data is generated when using the platform. For example, when a user uses a video platform, their behavior data includes, but is not limited to, watching a certain type of video or interacting with a video. Preferably, users on the video platform may actively initiate live streaming transmission as content producers, and their behavior data can include the status behavior data corresponding to the live streaming room, including but not limited to the status of viewers in the live streaming room and the interaction status of viewers in the live streaming room. This behavior data can extract necessary key feature data, providing reliable guidance for subsequent content recommendations tailored to the user.
[0051] Understandably, the collected behavioral data carries corresponding user tagging information, enabling subsequent analysis of the corresponding users and storage in the corresponding memory shards, thus avoiding data processing confusion and errors.
[0052] Preferably, user behavior data is collected synchronously through the backend data management center, which can ensure the complete collection of behavior data without hindering users from using the platform.
[0053] S300: Extract feature data from the behavioral data, wherein the feature data carries the user tagging information; In this embodiment, the behavioral data records a large amount of user transaction information in detail. It is necessary to extract key feature data from the behavioral data, delete daily operation data that has no reference value, and aggregate or process data that has research reference value, so as to make the user's operation characteristics more obvious and improve the efficiency of subsequent analysis.
[0054] Preferably, in this embodiment, a User Experience Design (UED) module is introduced to extract feature data. This UED module includes a data cleaning operator and several feature writing operators. The data cleaning operator filters out abnormal data from the input behavioral data and inputs the filtered behavioral data into the feature writing operators to complete the feature data extraction. It is understood that several feature writing operators can simultaneously extract feature data from multiple behavioral data sets, thereby improving extraction efficiency. Furthermore, when the UED module extracts feature data, it applies a Directed Acyclic Graph (DAG) to run and extract the data, which improves the flexibility and concurrency of the extraction process.
[0055] Specifically, such as Figure 3 As shown, the feature data extracted from the behavioral data includes: S310. Preset different user types and set the characteristic indicators corresponding to the user types; In this embodiment, because users of the platform have different attributes, they can generally include users who read or watch content and users who are content producers. The required feature perspectives for analysis differ, and therefore the feature data to be extracted will also differ. Therefore, it is necessary to preset different user types and set corresponding feature indicators for different user types to facilitate subsequent extraction. Preferably, user types can include viewing users and content-producing users. Feature indicators set for viewing users include, but are not limited to, the number of times various types of content are viewed, the viewing duration, and the interaction status; feature indicators set for content-producing users include, but are not limited to, the number of viewers, traffic, and interaction status of the content produced. Preferably, the content can include published media content such as videos, articles, and images.
[0056] S320. Determine the user type corresponding to the behavior data based on the user tag information carried by the behavior data; In this embodiment, behavioral data carries user tagging information, which reflects the user's operational attributes. For example, when a user uses a video platform, based on the user tagging information corresponding to behavioral data such as watching a certain type of video or interacting with a video, the user type can be determined to be a viewer; if a user actively initiates live streaming media transmission as a content producer, based on the user tagging information they carry, the user type can be determined to be a content producer.
[0057] S330. Determine the corresponding feature indicators according to the determined user type; S340. Based on the determined feature indicators, extract the feature data corresponding to the behavioral data.
[0058] In this embodiment, based on the determined feature indicators, feature data related to the feature indicators needs to be extracted from the behavioral data. For example, if the feature indicator is the number of times a game video is viewed, then two data points related to the feature indicator are obtained from the behavioral data: "Game A video was viewed 1 time; Game B video was viewed 2 times". The extracted feature data can be "Number of times a game video was viewed: 3".
[0059] S400: Based on the user tagging information of the feature data and the memory address information of the memory slice, write the feature data into the corresponding memory slice; In this embodiment, each user is allocated a corresponding memory slice. After obtaining the user's feature data, the feature data needs to be written to the corresponding memory slice. The corresponding memory slice needs to perform location based on the user tag information of the feature data and the memory address information of the memory slice.
[0060] Specifically, such as Figure 4As shown, the step of writing the feature data into the corresponding memory slice based on the user tag information of the feature data and the memory address information of the memory slice includes: S411. Construct a write request for the feature data based on the feature data; In this embodiment, constructing a write request can provide a write instruction or write task. After responding to the write request in the background, the writing of the feature data is completed. Preferably, this embodiment also includes a proxy service component for receiving the write request.
[0061] S412. Query the address mapping table according to the user tag information of the feature data, obtain the memory address information that matches the user tag information of the feature data, and determine the memory slice to be written according to the memory address information that matches the user tag information of the feature data. In this embodiment, by traversing the address mapping table, it is necessary to find the memory address information that matches the user tag information of the feature data. This allows us to find the corresponding memory slice and its memory location, thereby determining the memory slice to be written. Preferably, the traversal and determination are performed in the main service component.
[0062] Preferably, because the user tag information and the address mapping table may have format differences, directly querying the address mapping table with the obtained user tag information would consume a lot of comparison time. Therefore, the user tag information can be input into a hash function to obtain user tag information with the same format as the user information in the address mapping table, thereby enabling the matching memory address information to be found more easily and improving the query efficiency. Preferably, the format conversion is performed in the proxy service component.
[0063] S413. Forward the write request to the memory shard to be written, and complete the write operation of the feature data in the memory shard to be written.
[0064] In this embodiment, the write request is forwarded to the memory shard to be written, thereby realizing the writing of feature data. Preferably, the forwarding is performed by the proxy service component.
[0065] Specifically, the memory shard comprises several blocks, and the feature data written to the memory shard is distributed and stored in the memory shard in a distributed manner. The occurrence time of the stored feature data is synchronously written to each block.
[0066] In this embodiment, each memory slice is a contiguous storage space. A memory slice can be divided into several blocks, so that the feature data at each moment is stored in one block, and the occurrence time corresponding to that feature data is also stored, which facilitates subsequent tracing and retrieval of the feature data. It can be understood that in a memory slice, the occurrence times of the feature data corresponding to each adjacent block are continuous within a preset time interval. For example, if the preset time interval is 1 minute, when the occurrence time of the first block is 00:01, the adjacent second block should store the feature data with the occurrence time of 00:02.
[0067] Specifically, such as Figure 5 As shown, writing the feature data into the corresponding memory slice includes: S421. According to the occurrence time of the feature data, write the feature data and the occurrence time of the feature data into the corresponding memory slice block; In this embodiment, each memory slice is equipped with a pointer to identify the position of the header block and the position of the block where the previous storage operation was completed. When acquiring feature data, it is necessary to check whether the occurrence time corresponding to the block position where the previous storage operation was completed is contiguous with the occurrence time of the feature data to be written within a preset time interval, and whether the occurrence time corresponding to the block position where the previous storage operation was completed is earlier than the occurrence time of the feature data to be written. If both conditions are met, the feature data is written to the corresponding memory slice.
[0068] S422. If the feature data of the current occurrence moment is written to the tail block of the memory slice, then the feature data of the next occurrence moment is written to the head block of the memory slice, overwriting the data already stored in the head block.
[0069] In this embodiment, when the feature data of the next occurrence time is being written, if the block location marked by the relevant pointer indicating the completion of the previous storage operation is found to be the tail block of the memory slice, then the feature data of the next occurrence time needs to be written to the head block of the memory slice, overwriting the data already stored in the head block. The data already stored in the head block also includes the feature data and the corresponding occurrence time. Specifically, the position of the head block of the memory slice can be obtained based on the identifier of the relevant pointer, and the writing operation can be performed. By cyclically utilizing memory slices, the risk of infinite expansion of memory slices can be completely eliminated, invalid historical data accumulation can be eliminated, and the utilization rate of memory slices can be improved.
[0070] Specifically, in this embodiment, it also includes: Several fragment copies are pre-configured for the memory fragments; After writing the feature data into the corresponding memory shard, the written feature data is also synchronized to the shard copy corresponding to the memory shard. In this embodiment, when configuring memory shards, several shard replicas are configured for each memory shard. Specifically, each shard replica is a clone of the corresponding memory shard. Each shard replica has the same structure and stores the same data as the corresponding memory shard. When the feature data in the memory shard changes, the feature data in the corresponding memory replica is also updated synchronously. Therefore, in the event of a single point of failure in a memory shard, feature data can be read based on the corresponding shard replica, improving the disaster recovery capability for feature data processing. It is understood that after feature data is written to the corresponding memory shard, a consistency synchronization algorithm is needed to update the data of all shard replicas corresponding to that memory shard, ensuring that the data in all shard replicas of that memory shard is identical to that memory shard, preventing inaccurate data reads during operations. Subsequent read operations can be performed on the memory shard or its replicas, avoiding all read operations being performed on the same memory shard, thus improving data read efficiency through the shard replica configuration.
[0071] S500: Obtain guidance information for the data to be read, wherein the guidance information includes user tagging information; In this embodiment, when feature data needs to be read, guidance information for the corresponding data to be read is required to indicate which user's feature data was used to calculate the read data. Therefore, this guidance information needs to include the corresponding user guidance information. It is understood that the data to be read can be a combination of feature data within a preset time range corresponding to a certain feature indicator. Therefore, it is necessary to accurately obtain the corresponding feature data for combination to obtain accurate read data and improve the efficiency of subsequent analysis of the read data.
[0072] S600: Based on the user tag information of the guidance information and the memory address information of the memory slice, obtain the data to be read from the corresponding memory slice.
[0073] In this embodiment, after obtaining the guidance information for reading the data, it is necessary to locate the corresponding memory slice to read the data. The location of the corresponding memory slice is determined based on the user tag information in the guidance information and the memory address information of the memory slice.
[0074] Specifically, such as Figure 6 As shown, the process of obtaining the data to be read from the corresponding memory segment based on the user tag information of the guidance information and the memory address information of the memory segment includes: S611. Construct a read request for the data to be read based on the guidance information; In this embodiment, constructing a read request provides a read instruction or read task. After responding to the read request in the background, the read of the feature data is completed. Preferably, the read request is received through a proxy service component.
[0075] S612. Query the address mapping table according to the user tag information of the guidance information to obtain the memory address information that matches the user tag information of the guidance information, and determine the memory slice to be read according to the memory address information that matches the user tag information of the guidance information. In this embodiment, by traversing the address mapping table, it is necessary to find the memory address information that matches the user tag information of the guidance information. This allows the corresponding memory slice and its memory location to be found, thereby determining the memory slice to be read. Preferably, the traversal and determination are performed in the main service component.
[0076] Preferably, because the user tag information and the address mapping table may have format differences, directly querying the address mapping table with the obtained user tag information would consume a lot of comparison time. Therefore, the user tag information can be input into a hash function to obtain user tag information with the same format as the user information in the address mapping table, thereby enabling the matching memory address information to be found more easily and improving the query efficiency. Preferably, the format conversion is performed in the proxy service component.
[0077] S613. Forward the read request to the memory segment to be read, and complete the read operation of the data to be read in the memory segment to be read.
[0078] In this embodiment, the read request is forwarded to the memory shard to be read, thereby enabling the reading of the data to be read. Preferably, the forwarding is performed by the proxy service component.
[0079] Specifically, the guidance information for the data to be read also includes feature indicators and time range; In this embodiment, the memory slices contain various feature data. During subsequent analysis, it may be necessary to acquire a specific feature data for analysis. Therefore, the data to be read needs to include feature indicators, specifying the type of feature data to be read. Analyzing feature data at a single moment may not reveal the corresponding pattern; therefore, it is necessary to define a time range and obtain the sum of feature data over a certain period to better analyze the corresponding pattern. Thus, the guidance information needs to include a time range. Preferably, this time range includes the start and end times of reading the feature data.
[0080] like Figure 7As shown, obtaining the data to be read from the corresponding memory slice includes: S621. Based on the time range, determine the block to be read in the corresponding memory slice; In this embodiment, the time range includes the start and end times of reading feature data, and the block to be read can be determined based on the start and end times. For example, if a memory slice contains ten blocks, the preset time interval is one minute, and feature data occurring between 10:00 and 10:10 is stored, and the start time of the guidance information's time range is set to 10:06 and the end time to 10:10, then the block to be read can be determined to be the block corresponding to the occurrence times between 10:06 and 10:10. Preferably, this embodiment also includes a flexibly sized sliding window, and the size of the sliding window is set by the start and end times to quickly and clearly determine the block to be read. For example, in the above scenario, the sliding window is set to a length of five blocks, and according to the corresponding pointer, the head of the sliding window is placed in the block of the memory slice where the occurrence time is 10:06, and the tail of the sliding window is placed in the block of the memory slice where the occurrence time is 10:10.
[0081] S622. Based on the feature indicators of the guidance information, calculate the feature data of the block to be read to obtain the data to be read.
[0082] In this embodiment, the blocks to be read have been determined. Based on the feature indicators of the guidance information, the corresponding feature data in each block are summed to obtain the data to be read. For example, if the feature indicator to be read is the number of times a user watches a game video, and the time range is 10:06 to 10:10, then the feature data read based on the sliding window is as follows: Therefore, by summing the feature data for this time range, the data to be read is that the user watched 4 game videos between 10:06 and 10:10.
[0083] Specifically, based on the user tagging information of the guidance information and the memory address information of the memory slice, the data to be read is obtained from the corresponding memory slice, including: Based on the user tagging information of the guidance information and the memory address information of the memory shard, the data to be read is obtained from the corresponding memory shard and / or the corresponding shard copy.
[0084] In this embodiment, the reading operation can be performed from the corresponding memory shard or from a shard copy of the memory shard. This allows data to be read from the shard copy even if a single point of failure occurs in the memory shard, improving the disaster recovery capability for feature data processing. Furthermore, the reading task can be distributed and allocated to the shards, reducing the load on the memory shards and thus reducing the risk of memory shard failure.
[0085] like Figure 8 As shown in the embodiments of this application, a data processing system is also provided. Optionally, the system includes: The system includes a memory allocation module 711, a collection module 712, an extraction module 713, a writing module 714, a guidance information acquisition module 715, and a reading module 716, among which: The memory slice allocation module 711 is used to pre-allocate memory slices for users, wherein the memory slices are marked with memory address information; In this embodiment, the memory sharding allocation module 711 can be used to perform... Figure 2 For a detailed description of the memory sharding allocation module 711, please refer to the description of step S100 shown.
[0086] The acquisition module 712 is used to collect behavioral data carrying user tagging information; In this embodiment, the acquisition module 712 can be used to perform... Figure 2 For a detailed description of the acquisition module 712, please refer to the description of step S200 shown.
[0087] Extraction module 713 is used to extract feature data from the behavioral data, wherein the feature data carries the user tagging information; In this embodiment, the extraction module 713 can be used to perform... Figure 2 For a detailed description of the extraction module 713, please refer to the description of step S300 shown.
[0088] The writing module 714 is used to write the feature data into the corresponding memory segment based on the user tag information of the feature data and the memory address information of the memory segment; In this embodiment, the writing module 714 can be used to perform... Figure 2 For a detailed description of the writing module 714, please refer to the description of step S400 shown.
[0089] The guidance information acquisition module 715 is used to acquire guidance information of the data to be read, wherein the guidance information includes user tag information; In this embodiment, the guidance information acquisition module 715 can be used to perform... Figure 2 For a detailed description of the guidance information acquisition module 715, please refer to the description of step S500 shown.
[0090] The reading module 716 is used to obtain the data to be read from the corresponding memory segment based on the user tag information of the guidance information and the memory address information of the memory segment.
[0091] In this embodiment, the reading module 716 can be used to perform... Figure 2 For a detailed description of the reading module 716, please refer to the description of step S600 shown.
[0092] This application also provides an electronic device, the structure of which is as follows: Figure 9 As shown, the electronic device includes a memory 811, a processor 812, a communication module 813, and an input / output interface 814, etc. Optionally, the memory 811, the processor 812, the communication module 813, and the input / output interface 814 can be connected and communicate with each other through a bus 815.
[0093] The memory 811 is used to store one or more computer programs and to transfer the code of the computer programs to the processor 812; when the one or more computer programs are executed by the processor 812, a data processing method in this embodiment of the application is implemented.
[0094] Optionally, the electronic device can be connected to a network via communication module 813 to communicate with other devices, such as terminals or servers, and to interact with data. The electronic device can be various forms of digital computers, exemplarily such as desktop computers, servers, workbenches, mainframes, or other types of computers. The electronic device can also be various forms of mobile terminals, exemplarily such as smartphones, tablets, wearable devices (such as helmets, glasses, watches, etc.), and other similar mobile terminals.
[0095] Optionally, the electronic device can connect to desired input / output devices, such as a keyboard or display device, via the input / output interface 814. The electronic device itself may have a display device, and other display devices can also be connected externally via the input / output interface 814. Optionally, a storage device, such as a hard disk, can also be connected via the input / output interface 814 to store data from the electronic device, read data from the storage device, or store data from the storage device in the memory 811. It is understood that the input / output interface 814 can be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 814 can be a component of the electronic device or an external device connected to the electronic device when needed.
[0096] Optionally, the memory 811 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory, etc.
[0097] Optionally, the computer program stored in the memory 811 can be divided into one or more modules, which are stored in the memory 811 and executed by the processor 812 to perform the method provided in this embodiment. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device.
[0098] Optionally, the processor 812 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 812 include, but are not limited to, a central processing unit, a graphics processing unit, a digital signal processor, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, and can also be any suitable controller, microcontroller, processor, etc. The processor 812 executes the various methods and processes of this embodiment, exemplarily, such as a data processing method according to an embodiment of this application.
[0099] Optionally, the bus 815 may include a path for transmitting information. Depending on its function, the bus 815 may be divided into an address bus, a data bus, a control bus, etc.
[0100] In an optional implementation, this application embodiment also provides a computer storage medium storing a computer program thereon. When the computer program is executed by a computer, it enables the computer to perform the methods described in the above-described method embodiments. Part or all of the computer program can be loaded and / or installed on the memory 811 of an electronic device. When the computer program is executed by the processor 812, one or more steps of a data processing method according to an embodiment of this application can be performed.
[0101] Optionally, the computer-readable storage medium may be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.
[0102] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention should be included within the protection scope of the claims of the present invention.
Claims
1. A data processing method, characterized in that, The method includes: Memory slices are pre-allocated to users, and the memory slices are marked with memory address information; Collect behavioral data carrying user tagging information; Extract feature data from the behavioral data, wherein the feature data carries the user tagging information; Based on the user tagging information of the feature data and the memory address information of the memory slice, the feature data is written into the corresponding memory slice; Obtain guidance information for the data to be read, wherein the guidance information includes user tagging information; Based on the user tag information of the guidance information and the memory address information of the memory slice, the data to be read is obtained from the corresponding memory slice.
2. The method according to claim 1, characterized in that, The method further includes: Establish an address mapping table for user and memory address information for the memory slices; The process of writing the feature data into the corresponding memory slice based on the user tagging information of the feature data and the memory address information of the memory slice includes: Construct a write request for the feature data based on the feature data; The address mapping table is queried according to the user tag information of the feature data to obtain the memory address information that matches the user tag information of the feature data, and the memory slice to be written is determined according to the memory address information that matches the user tag information of the feature data. The write request is forwarded to the memory shard to be written, and the write operation of the feature data is completed in the memory shard to be written.
3. The method according to claim 1, characterized in that, The method further includes: Establish an address mapping table for user and memory address information for the memory slices; The process of retrieving the data to be read from the corresponding memory segment based on the user tagging information of the guidance information and the memory address information of the memory segment includes: Construct a read request for the data to be read based on the guidance information; The address mapping table is queried according to the user tag information of the guidance information to obtain the memory address information that matches the user tag information of the guidance information, and the memory slice to be read is determined according to the memory address information that matches the user tag information of the guidance information. The read request is forwarded to the memory shard to be read, and the read operation of the data to be read is completed in the memory shard to be read.
4. The method according to claim 1, characterized in that, The feature data extracted from the behavioral data includes: Different user types are preset, and corresponding feature indicators are set for each user type; Based on the user tagging information carried by the behavioral data, the user type corresponding to the behavioral data is determined; Determine the corresponding feature indicators based on the identified user type; Based on the determined feature indicators, extract the feature data corresponding to the behavioral data.
5. The method according to claim 1, characterized in that, The method further includes: Several fragment copies are pre-configured for the memory fragments; After writing the feature data into the corresponding memory shard, the written feature data is also synchronized to the shard copy corresponding to the memory shard. Based on the user tagging information of the guidance information and the memory address information of the memory slice, the data to be read is obtained from the corresponding memory slice, including: Based on the user tagging information of the guidance information and the memory address information of the memory shard, the data to be read is obtained from the corresponding memory shard and / or the corresponding shard copy.
6. The method according to any one of claims 1 to 5, characterized in that, The memory shard comprises several blocks, and the feature data written to the memory shard is distributed and stored in the memory shard. The occurrence time of the stored feature data is synchronously written to each block.
7. The method according to claim 6, characterized in that, The step of writing the feature data into the corresponding memory slice includes: According to the occurrence time of the feature data, the feature data and the occurrence time of the feature data are written into the corresponding memory slice block; If the feature data of the current occurrence is written to the tail block of the memory slice, then the feature data of the next occurrence is written to the head block of the memory slice, overwriting the data already stored in the head block.
8. The method according to claim 6, characterized in that, The guidance information for the data to be read also includes feature indicators and time range; The step of obtaining the data to be read from the corresponding memory slice includes: Based on the time range, determine the block to be read in the corresponding memory slice; Based on the feature indicators of the guidance information, the feature data in the block to be read is calculated to obtain the data to be read.
9. A data processing system, characterized in that, The system includes: The memory shard allocation module is used to pre-allocate memory shards for users, and the memory shards are marked with memory address information; The data collection module is used to collect behavioral data carrying user tagging information; The extraction module is used to extract feature data from the behavioral data, wherein the feature data carries the user tagging information; The writing module is used to write the feature data into the corresponding memory segment based on the user tag information of the feature data and the memory address information of the memory segment; The guidance information acquisition module is used to acquire guidance information for the data to be read, wherein the guidance information includes user tag information; The reading module is used to obtain the data to be read from the corresponding memory segment based on the user tag information of the guidance information and the memory address information of the memory segment.
10. An electronic device, characterized in that, include: Memory, used to store one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements a data processing method as described in any one of claims 1-8.