Data micro-batching and table storage method and system
Through the data micro-batched table storage method, local cache memory queues and worker thread pools are used for batch processing and distribution, the problem of low write efficiency of buried point data is solved, low latency and high performance data table is realized, and database pressure is reduced.
Patent Information
- Application Number
- CN202510279407.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-11
AI Technical Summary
In the prior art, the low efficiency of writing buried point data will lead to increased database pressure, degradation of performance, and may even lead to downtime.
The data microbatch table storage method is adopted to batch process the target data through a local cache memory queue, and is distributed to the worker thread pool for processing through a scheduling thread for processing and storage in the database. When the local cache memory queue is full, cache compensation processing is performed to prevent data loss.
Through microbatch processing, low-latency and high-performance data tables are achieved, which meets the real-time requirements of data tables, and reduces the concurrency and pressure on database operations.
Smart Images

Figure CN119781992B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data technology, relates to a data processing technology, and particularly relates to a method and system for storing data in batches and falling tables in a micro-batch manner. Background Art
[0002] Generally, in order to provide users with a better APP experience and more accurate services, it is often necessary to collect user behavior buried point data. These data help to understand the deficiencies of the product, improve the user portrait, support operation decisions, and calculate core indicators to evaluate the operation status and market performance of the product.
[0003] When users browse and operate the app, such as opening pages, clicking on page elements, swiping, etc., corresponding behavior event buried point data will be generated. The client app reports the collected buried point data to the server through an interface for data analysis. Conventional buried point data processing is implemented through hard coding, which is not flexible enough. Any logical change requires developing code, with a complex process, high labor costs, and slow iteration speed; the database write pressure is increasing day by day, and the database has performance bottlenecks.
[0004] When the number of users is increasing day by day, the buried point data generated by the client is becoming more and more huge, the requests for the client to upload buried point data are increasing, and the pressure on the server and the database is becoming greater and greater.
[0005] The pressure on the server can generally be relieved by horizontal expansion, deploying multiple server nodes to share the pressure. Considering cost and architectural complexity, in general, most companies adopt a master-slave database architecture. The traffic of data writing cannot be horizontally expanded, and the master data bears a very high concurrency, which will cause a sudden increase in the pressure on the master database, resulting in a decline in database performance, slow response, occupation and blockage of connection numbers, etc. In severe cases, the database may even crash. The main reasons are the iteration of buried point data, resulting in complex and variable requirements for buried point data processing; and in the case of a very large amount of buried point data, the database write concurrency is very high, resulting in a decrease in data write efficiency. Summary of the Invention
[0006] The purpose of this application is to provide a method and system for storing data in batches and falling tables in a micro-batch manner to solve the problem of low efficiency of writing buried point data in the prior art.
[0007] In a first aspect, this application provides a method for storing data in batches and falling tables in a micro-batch manner, and the method includes:
[0008] Obtain target data to be processed, and determine whether the target data meets a preset condition;
[0009] After determining that the target data meets the preset condition, send the target data to the local cache memory queue, and determine whether the local cache memory queue is full;
[0010] After determining that the local cache memory queue is not full, store the target data in the local cache memory queue;
[0011] According to the maximum number of tasks pulled from the local cache memory queue at one time, the scheduling thread regularly pulls the data to be consumed to be stored from the local cache memory queue, and allocates the data to be consumed to each working thread according to the parameters of the working thread pool, and stores the data to be consumed in the database after being processed by the working thread;
[0012] After determining that the local cache memory queue is full, perform cache compensation processing on the target data.
[0013] In one implementation manner of the first aspect, the performing cache compensation processing on the target data includes:
[0014] Judge whether the target data fails to be sent to the local cache memory queue for the first time;
[0015] If the target data fails to be sent to the local cache memory queue for the first time, send scheduling information to the scheduling thread, and the scheduling thread immediately pulls data from the local cache memory queue according to the maximum number of tasks pulled from the local cache memory queue at one time and sends it to the working thread pool;
[0016] Resend the target data to the local cache memory queue, and re-judge whether the local cache memory queue is full.
[0017] In one implementation manner of the first aspect, if the target data does not fail to be sent to the local cache memory queue for the first time, store the target data in the MQ distributed queue, and after a preset time, send the target data to the failure compensation processing program, and the failure compensation processing program resubmits and sends the target data to the local cache memory queue.
[0018] In one implementation manner of the first aspect, the working thread pool includes core working threads and standby working threads. When all the core working threads in the working thread pool are occupied, the standby working threads are enabled to process data, and when all the core working threads and the standby working threads in the working thread pool are occupied, the working thread pool stores the remaining data scheduled by the scheduling thread in the task warehouse.
[0019] In one implementation manner of the first aspect, the allocating the data to be consumed to each working thread according to the parameters of the working thread pool includes:
[0020] Obtain the number of unoccupied worker threads in the worker thread pool, and allocate data to the worker threads according to the single-work capacity of each unoccupied worker thread;
[0021] If there is remaining data after the allocation of the data to be consumed, store the remaining data in the task warehouse, and submit the remaining data to the worker threads that have completed data processing in the order of the time when each worker thread completes data processing;
[0022] Among them, the quantity of the remaining data submitted to each worker thread that has completed data processing each time does not exceed the single-work capacity of the worker thread.
[0023] In one implementation manner of the first aspect, the preset conditions include at least one of a screening condition and a default rule, and the screening condition includes at least one of a page name, an event type, and an event name.
[0024] In one implementation manner of the first aspect, the method further includes:
[0025] If the target data does not meet the preset conditions, abandon storing the target data, and continue to determine whether the next target data meets the preset conditions.
[0026] In one implementation manner of the first aspect, the target data includes at least any one of buried point data and device information data. When the target data is buried point data, the method further includes: preprocessing the buried point data to standardize the format and content of the buried point data.
[0027] In one implementation manner of the first aspect, the preprocessing of the buried point data includes:
[0028] After the user generates buried point data, the buried point data is reported to the server regularly through the client;
[0029] After the server receives the buried point data, write the buried point data into the message middleware, and asynchronously read the buried point data from the message middleware;
[0030] Perform parsing processing on the read buried point data to obtain buried point data in a standard format, and use the buried point data in the standard format as the target data.
[0031] In a second aspect, the present invention provides a data micro-batch down-table storage system, and the system includes:
[0032] A judgment module, configured to obtain target data to be processed and judge whether the target data meets preset conditions;
[0033] A sending module, configured to send the target data to a local cache memory queue after determining that the target data meets a preset condition, and determine whether the local cache memory queue is full;
[0034] A first processing module, configured to store the target data into the local cache memory queue after determining that the local cache memory queue is not full;
[0035] A scheduling module, configured to regularly pull the data to be consumed to be stored from the local cache memory queue by a scheduling thread according to the maximum number of tasks pulled by the local cache memory queue at a single time, and allocate the data to be consumed to each working thread according to the parameters of the working thread pool, and store the data to be consumed in a database after being processed by the working thread;
[0036] A second processing module, configured to perform cache compensation processing on the target data after determining that the local cache memory queue is full.
[0037] As described above, the data micro-batching and table-drop storage method and system of the present application have the following beneficial effects:
[0038] In the present application, the target data to be stored is batch-saved by adopting a micro-batching processing method. The target data is sequentially placed into the local cache memory queue, and the data in the local cache memory queue is regularly scheduled by a scheduling thread and distributed to the working threads in the working thread pool for processing. After being processed by the working thread, the scheduled data is sequentially stored in the database. When the local cache memory queue is full, cache compensation processing is performed on the target data to ensure that the target data can be cached in the distributed queue after the second submission fails, prevent data loss, and re-submit the target data to the local cache memory queue after a preset time. Through the ingenious combination of micro and batch, it has the characteristics of low latency and high performance, which not only meets the real-time requirements of data table dropping, but also reduces the concurrency and pressure on database operations. Moreover, the method of the present application has high generality and reusability and can be applied to the data table dropping processing of different task business scenarios. Description of the Drawings
[0039] Figure 1 It shows a flowchart of the data micro-batching and table-drop storage method described in the embodiment of the present application.
[0040] Figure 2 It shows a schematic diagram of the execution process of the data micro-batching and table-drop storage method described in the embodiment of the present application.
[0041] Figure 3 It shows a schematic diagram of the process of preprocessing the buried point data in the data micro-batching and table-drop storage method described in the embodiment of the present application.
[0042] Figure 4 It shows a structural block diagram of the data micro-batch table-drop storage system described in the embodiments of the present application. Specific embodiments
[0043] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0044] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0045] Refer to Figures 1 to 3 , the following embodiments of the present application provide the data micro-batch table-drop storage method and system described in the present application, and have the following beneficial effects:
[0046] The data micro-batch table-drop storage method of the present application performs batch saving processing on the target data to be stored by adopting the micro-batch processing method. By sequentially putting the target data into the local cache memory queue, and scheduling the data in the local cache memory queue regularly by a scheduling thread and distributing it to the working threads in the working thread pool for processing. After the working threads process, the scheduled data is sequentially stored in the database. When the local cache memory queue is full, cache compensation processing is performed on the target data to ensure that it can be cached in the distributed queue after the second submission of the target data fails, preventing data loss, and resubmitting the target data to the local cache memory queue after reaching the preset time. Through the ingenious combination of micro and batch, it has the characteristics of low latency and high performance, meeting both the real-time requirements of data table-drop and reducing the concurrency and pressure on database operations. Moreover, the method of the present application scheme has high generality and reusability, and can be applied to data table-drop processing in different task business scenarios.
[0047] Next, the technical solutions in the embodiments of the present application will be described in detail with reference to the accompanying drawings in the embodiments of the present application.
[0048] As Figure 1 shown, this embodiment provides a data micro-batch table-drop storage method, and the method includes the following steps:
[0049] S101. Obtain target data to be processed, and determine whether the target data meets a preset condition.
[0050] In this embodiment, for the target data to be processed, first determine whether the target data meets a preset condition. Considering that not all data needs to be stored in the database by falling into a table, a preset condition is set to preliminarily judge the target data, ensuring that only the target data that meets the preset condition will be stored in the database, reducing the storage pressure on the database.
[0051] In some embodiments, the preset condition includes at least one of a screening condition and a default rule, and the screening condition includes at least one of a page name, an event type, and an event name.
[0052] Specifically, the preset condition includes at least one of a screening condition and a default rule. The screening condition is a whitelist rule, and the default rule is a blacklist rule. The two can be configured according to actual situations. For example, the screening condition includes at least one of a page name, an event type, and an event name. By judging whether the target data falls into the screening condition, the target data is screened in the form of a whitelist, facilitating the processing of the target data.
[0053] Exemplarily, for example, the page name includes "borrow cash - pre - loan". When there is data in the target data whose page name contains "borrow cash - pre - loan", it is determined that the target data meets the preset condition; otherwise, it does not.
[0054] S102. After determining that the target data meets the preset condition, send the target data to the local cache memory queue, and determine whether the local cache memory queue is full.
[0055] In this embodiment, refer to Figure 2 , after determining that the target data meets the preset condition, it can be determined that the target data needs to be stored in the database by falling into a table. Then, the target data is sent to the local cache memory queue, and it is determined whether the data in the local cache memory queue is full. Since a large amount of data is stored in the local cache memory queue, in order to avoid the situation where the target data to be stored has no storage location and cannot be stored, by determining whether the local cache memory queue is full, different processing is performed on the target data, ensuring that the target data can be added to the local cache memory queue in time when the local cache memory queue is not full, and can also be processed in time when the local cache memory queue is full, preventing data loss.
[0056] On the other hand, the method further includes:
[0057] If the target data does not meet the preset conditions, discard the storage of the target data and continue to determine whether the next target data meets the preset conditions.
[0058] In this embodiment, after determining that the target data does not meet the preset conditions, it is determined that the current target data does not need to be stored, so it can be directly discarded, and the next target data is processed continuously to determine whether the next target data meets the preset conditions, so as to realize the continuous processing of the target data.
[0059] S103. After determining that the local cache memory queue is not full, store the target data in the local cache memory queue.
[0060] S104. According to the maximum number of tasks pulled from the local cache memory queue at one time, the scheduling thread periodically pulls the data to be consumed to be stored from the local cache memory queue, and distributes the data to be consumed to each working thread according to the parameters of the working thread pool, and stores the data to be consumed in the database after being processed by the working thread.
[0061] In this embodiment, after determining that the local cache memory queue is not full, the target data is directly submitted and stored in the local cache memory queue. Since the scheduling thread periodically pulls the data to be consumed stored in the local cache memory queue according to the preset time, and distributes the pulled data to be consumed to each working thread in the working thread pool, so that the working thread processes the distributed data to be consumed and then stores it in the database.
[0062] In some other embodiments, the working thread pool includes core working threads and standby working threads. After all the core working threads in the working thread pool are occupied, the standby working threads are enabled to process the data, and when all the core working threads and the standby working threads in the working thread pool are occupied, the working thread pool stores the remaining data scheduled by the scheduling thread in the task warehouse.
[0063] In this embodiment, the working thread pool includes core working threads and standby working threads. The core working threads are used as default threads to process and save the data to the database, while the standby working threads are used as standby threads. After all the core working threads are occupied, the standby working threads are used to process the remaining data and send it to the database. When all the core working threads and the standby working threads are occupied and there is still remaining data to be consumed scheduled by the scheduling thread, the remaining data is stored in the task warehouse in the scheduling thread. After the core working threads or the standby working threads complete the task processing, the remaining data is continued to be distributed, thus effectively preventing the situation of accumulation in the working thread pool.
[0064] In some further embodiments, allocating the data to be consumed to each worker thread according to the parameters of the worker thread pool includes:
[0065] Obtaining the number of unoccupied worker threads in the worker thread pool, and allocating data to the worker threads according to the single-work capacity of each unoccupied worker thread;
[0066] If there is remaining data after the allocation of the data to be consumed, storing the remaining data in the task warehouse, and submitting the remaining data to the worker threads that have completed data processing in the order of the time when each worker thread completes data processing;
[0067] Wherein, the quantity of the remaining data submitted to each worker thread that has completed data processing each time does not exceed the single-work capacity of the worker thread.
[0068] In this embodiment, after the scheduling thread invokes the data to be consumed from the local cache memory queue, first obtain the number of unoccupied worker threads in the worker thread pool, including core worker threads and standby worker threads, and the allocation priority of the core worker threads is higher than that of the standby worker threads. Allocate data to each worker thread according to the single-work capacity of each unoccupied worker thread, so that the worker threads can process the data and save it to the database.
[0069] During the process of allocating data, when there is still remaining data after the allocation of the data to be consumed, the remaining data is stored in the task warehouse in the worker thread pool. Subsequently, as the worker threads complete data processing in chronological order, the remaining data is submitted to the worker threads that have completed data processing in sequence, ensuring that when the capacity of the data to be consumed is greater than the current processing capacity of the worker thread pool, there will be no data accumulation.
[0070] It should be noted that the capacity of the remaining data stored in the task warehouse does not exceed the capacity upper limit of the task warehouse. For data exceeding the capacity upper limit of the task warehouse, it waits to be put in after the data stored in the task warehouse is processed or after the worker threads complete data processing.
[0071] Exemplarily, the maximum capacity of the local cache memory queue is 10,000, the time interval for the scheduling thread to schedule regularly is 50 ms, the maximum number of data pulled by the scheduling thread from the local cache memory queue each time is 2,000, the maximum number of data saved in batches to the database by each working thread in the working thread pool is 100, the maximum number of core working threads in the working thread pool is 10, the number of standby working threads is 10, and the maximum capacity of the task repository in the working thread pool is 200.
[0072] S105. After determining that the local cache memory queue is full, perform cache compensation processing on the target data.
[0073] After determining that the local cache memory queue is full, in order to prevent the loss of the target data, perform compensation processing on the target data, so that the target data can be submitted and stored again when the data in the local cache memory queue is not full later, ensuring the integrity of the target data and preventing data loss.
[0074] In some embodiments, the performing cache compensation processing on the target data includes:
[0075] Determine whether the target data fails to be sent to the local cache memory queue for the first time;
[0076] If the target data fails to be sent to the local cache memory queue for the first time, send a scheduling message to the scheduling thread, and the scheduling thread immediately pulls data from the local cache memory queue according to the maximum number of tasks pulled by the local cache memory queue each time and sends it to the working thread pool;
[0077] Resend the target data to the local cache memory queue, and re-determine whether the local cache memory queue is full.
[0078] In this embodiment, when a request for the target data is sent to the local cache memory queue and the local cache memory queue is full, in order to ensure that the target data will not be lost, first determine whether the current target data fails to be submitted to the local cache memory queue for the first time, so that when the local cache memory queue is still full after multiple submissions, the target data can be saved in the distributed queue for continued submission next time.
[0079] Specifically, after the local cache memory queue is full and the submission of target data fails, if it is determined that the target data fails to be submitted for the first time, a scheduling message is sent to the scheduling thread. After receiving the scheduling message, the scheduling thread immediately pulls data from the local cache memory queue according to the maximum number of tasks pulled from the local cache memory queue once, and sends the data to the worker thread pool, so as to reduce the data in the local cache memory queue. Then, the target data that failed to be submitted for the first time is sent to the local cache memory queue again, and it is determined again whether the local cache memory queue is full. After re-submission, if it is determined that the local cache memory queue is not full, the re-submitted target data is directly stored in the local cache memory queue in order. If it is determined that the local cache memory queue is still full, it is further determined whether the local cache memory queue is full.
[0080] Further, if the target data fails to be sent to the local cache memory queue for the first time, the target data is stored in the MQ distributed queue, and after a preset time is reached, the target data is sent to the failure compensation handler, and the failure compensation handler re-submits the target data and sends it to the local cache memory queue.
[0081] In this embodiment, when it is determined that the target data fails to be sent to the local cache memory queue for the first time. For example, after the initial submission fails, after the scheduling thread receives the scheduling message and schedules the local cache memory queue once immediately, if the target data that failed to be submitted is sent to the local cache memory queue again and the local cache memory queue is still full, it is determined that the current target data fails to be submitted for the second time. To avoid repeated submissions from affecting the efficiency of data storage in the table and prevent the loss of target data, after the target data fails to be submitted for the second time, the target data is cached in the distributed queue so that after a preset time is reached subsequently, the target data cached in the distributed queue is sent to the failure compensation handler, and the failure compensation handler re-sends the target data cached in the distributed queue to the local cache memory queue, ensuring that the target data that failed to be submitted can also be stored in the database through the local cache memory queue and written to the table.
[0082] It should be noted that the preset time is a data set manually, mainly to ensure that the failure compensation handler sends the target data when the local cache memory queue is not full, so as to ensure that the target data that failed to be submitted can be saved normally.
[0083] In some other embodiments, the target data includes at least any one of the buried point data and the device information data.
[0084] In the above method, the target data can be either buried point data or device information data, and the above method can be used to achieve batch storage in the table.
[0085] It should be noted that the target data can also be other types of data. This solution does not make special limitations on this. For different data, only the preset conditions need to be adjusted to achieve preliminary screening, and then the above method can be directly used for storage in the table, which has good reusability and can meet the requirements of storing different types of data in the table.
[0086] When the target data is buried point data, the method further includes: preprocessing the buried point data to standardize the format and content of the buried point data.
[0087] In some embodiments, the preprocessing of the buried point data refers to Figure 3 , including:
[0088] After the user generates buried point data, the client regularly reports the buried point data to the server;
[0089] After the server receives the buried point data, it writes the buried point data into the message middleware and asynchronously reads the buried point data from the message middleware;
[0090] Parse and process the read buried point data to obtain the buried point data in the standard format, and use the buried point data in the standard format as the target data.
[0091] In this embodiment, when the target data is buried point data, since the buried point data is becoming more and more huge with the increasing number of users, it is easy to cause greater pressure on the database. After the buried point data is collected in the solution of this application, the client regularly reports the buried point data to the server. After the server receives the buried point data, it first writes the buried point data into the message middleware kafka in sequence, and then asynchronously reads the buried point data from the message middleware kafka. Thus, the function of peak shaving and valley filling can be achieved according to the server resources.
[0092] Since there are differences between the buried point data, in order to further ensure the efficiency of storage in the table, the buried point data is parsed to obtain the buried point data in the standardized format, which is convenient for subsequent storage of the buried point data.
[0093] Specifically, since the buried point data reported by clients of different versions may vary, such as single reporting, batch reporting, inconsistent structures, inconsistent field naming, inconsistent field content code tables, buried points for dirty data, buried point default values, supplementary additional information, etc., buried point parsing is to convert these differentially structured buried point data into standard data formats and contents, facilitating subsequent analysis and use.
[0094] It should be noted that the parsing of buried point data in different formats to convert them into standard formats adopts the content of existing technologies. This solution does not involve improvements to the specific parsing process. Any technical means capable of converting buried point data into standard formats can be applied to the parsing process of this application's solution, and this solution does not make special limitations in this regard.
[0095] The data micro-batch landing table storage method of this application performs batch saving processing on the target data to be stored by adopting the method of micro-batch processing. By sequentially putting the target data into the local cache memory queue, and scheduling the data in the local cache memory queue regularly by a scheduling thread and distributing it to the working threads in the working thread pool for processing. After the working threads process it, the scheduled data is sequentially stored in the database. When the local cache memory queue is full, cache compensation processing is performed on the target data to ensure that it can be cached in the distributed queue after the second submission of the target data fails, preventing data loss, and resubmitting the target data to the local cache memory queue after a preset time. Through the ingenious combination of micro and batch, it has the characteristics of low latency and high performance, meeting both the real-time requirements of data landing tables and reducing the concurrency and pressure on database operations. Moreover, the method of this application's solution has high generality and reusability and can be applied to the data landing table processing of different task business scenarios.
[0096] The protection scope of the data micro-batch landing table storage method described in the embodiments of this application is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps and replacing steps of existing technologies according to the principles of this application is included in the protection scope of this application.
[0097] This application also provides a data micro-batch landing table storage system. Refer to Figure 4 , the system includes:
[0098] A judgment module 401, configured to obtain the target data to be processed and judge whether the target data meets a preset condition;
[0099] A sending module 402, configured to, after determining that the target data meets the preset condition, send the target data to the local cache memory queue and judge whether the local cache memory queue is full;
[0100] The first processing module 403 is configured to store the target data into the local cache memory queue after determining that the local cache memory queue is not full;
[0101] The scheduling module 404 is configured to pull the data to be consumed to be stored from the local cache memory queue regularly through a scheduling thread according to the maximum number of tasks pulled from the local cache memory queue at a time, and allocate the data to be consumed to each working thread according to the parameters of the working thread pool, and store the data to be consumed in the database after being processed by the working thread;
[0102] The second processing module 405 is configured to perform cache compensation processing on the target data after determining that the local cache memory queue is full.
[0103] The embodiment of the present application further provides a data micro-batch table storage system, which can implement the data micro-batch table storage method described in the present application. However, the implementation device of the data micro-batch table storage method described in the present application includes, but is not limited to, the structure of the data micro-batch table storage system listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principle of the present application are included in the protection scope of the present application.
[0104] It should be noted that the structures and principles of the above-mentioned modules correspond one by one to the steps in the above-mentioned data micro-batch table storage method, and the specific working principles can also refer to the introduction of the data micro-batch table storage method in the foregoing embodiments, so details are not described herein again.
[0105] The embodiment of the present invention further provides an electronic device, which includes: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the above-mentioned data micro-batch table storage method.
[0106] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by an electronic device, the above-mentioned data micro-batch table storage method is implemented.
[0107] Those of ordinary skill in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing a processor through a program. The program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)).
[0108] In several embodiments provided by the present application, it should be understood that the disclosed system, apparatus, or method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules or units can be in electrical, mechanical, or other forms.
[0109] The modules / units described as separate components may or may not be physically separated. The components shown as modules / units may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the objectives of the embodiments of the present application. For example, in each embodiment of the present application, the various functional modules / units can be integrated in a processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated in one module / unit.
[0110] Those of ordinary skill in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0111] The descriptions of the processes or structures corresponding to the above respective drawings each have their own emphasis. For parts not detailed in a certain process or structure, reference can be made to the relevant descriptions of other processes or structures.
[0112] The above embodiments are only illustrative of the principles and effects of this application and are not used to limit this application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in this application should still be covered by the claims of this application.
Claims
1. A data micro-batch table storage method, characterized in that: The method comprises: Obtaining target data to be processed, and determining whether the target data meets preset conditions; After determining that the target data meets the preset condition, sending the target data to the local cache memory queue, and determining whether the local cache memory queue is full; After determining that the local cache memory queue is not full, storing the target data in the local cache memory queue; According to the maximum number of tasks pulled at a time by the local cache memory queue, the data to be stored and consumed are pulled from the local cache memory queue by scheduling threads, and the data to be consumed are distributed to each worker thread according to the parameters of the worker thread pool, and the data to be consumed are stored in the database after being processed by the worker thread; After determining that the local cache memory queue is full, performing cache compensation processing on the target data; The performing cache compensation processing on the target data includes: Determining whether the target data fails to be sent to the local cache memory queue for the first time; If the target data fails to be sent to the local cache memory queue for the first time, a scheduling message is sent to the scheduling thread, and the scheduling thread immediately pulls data from the local cache memory queue according to the scheduling message according to the maximum number of tasks pulled at a time by the local cache memory queue and sends it to the worker thread pool; Re-sending the target data to the local cache memory queue, and re-determining whether the local cache memory queue is full; If the target data fails to be sent to the local cache memory queue for the first time, the target data will be stored in the MQ distributed queue, and after a preset time, the target data will be sent to the failure compensation handler, and the target data will be resubmitted and sent to the local cache memory queue through the failure compensation handler.
2. The data micro-batch table storage method according to claim 1 is characterized in that: The work thread pool includes core work threads and backup work threads. When the core work threads in the work thread pool are all occupied, the backup work threads are enabled to process data. When the core work threads and the backup work threads in the work thread pool are all occupied, the work thread pool stores the remaining data scheduled by the scheduling thread in the task warehouse.
3. The data micro-batch table storage method according to claim 2 is characterized in that: The allocating the to-be-consumed data to each worker thread according to the parameters of the worker thread pool includes: Obtaining the number of unoccupied working threads in the working thread pool, and allocating data to the working threads according to the single working capacity of each unoccupied working thread; If there is remaining data after the data to be consumed is allocated, the remaining data is stored in the task warehouse, and the remaining data is submitted to the working thread that has completed the data processing in sequence according to the time sequence in which each working thread completes the data processing; The amount of the remaining data submitted to the working thread that completes data processing each time does not exceed the single working capacity of the working thread.
4. The data micro-batch table storage method according to claim 1 is characterized in that: The preset condition includes at least one of a screening condition and a default rule, and the screening condition includes at least one of a page name, an event type, and an event name.
5. The data micro-batch table storage method according to claim 4 is characterized in that: The method further comprises: If the target data does not meet the preset condition, the target data is abandoned from being stored, and it is continued to be determined whether the next target data meets the preset condition.
6. The data micro-batch table storage method according to any one of claims 1 to 5, characterized in that: The target data includes at least any one of buried point data and device information data. When the target data is buried point data, the method further includes: preprocessing the buried point data to standardize the format and content of the buried point data.
7. The data micro-batch table storage method according to claim 6 is characterized in that: The preprocessing of the buried point data includes: After the user generates the buried point data, the buried point data is regularly reported to the server through the client; After receiving the buried data, the server writes the buried data into the message middleware, and asynchronously reads the buried data from the message middleware; The read buried point data is parsed to obtain the buried point data in a standard format, and the buried point data in the standard format is used as the target data.
8. A data micro-batch table storage system, characterized in that: The system comprises: A judgment module, used to obtain target data to be processed and judge whether the target data meets preset conditions; A sending module, configured to send the target data to a local cache memory queue after determining that the target data meets a preset condition, and determine whether the local cache memory queue is full; A first processing module, configured to store the target data into the local cache memory queue after determining that the local cache memory queue is not full; The scheduling module is used to pull the data to be stored and consumed from the local cache memory queue according to the maximum number of tasks pulled at a time by the local cache memory queue through the scheduling thread, and distribute the data to be consumed to each worker thread according to the parameters of the worker thread pool, and store the data to be consumed in the database after the worker thread processes it; A second processing module, configured to perform cache compensation processing on the target data after determining that the local cache memory queue is full; The step of performing cache compensation processing on the target data includes: Determining whether the target data fails to be sent to the local cache memory queue for the first time; If the target data fails to be sent to the local cache memory queue for the first time, a scheduling message is sent to the scheduling thread, and the scheduling thread immediately pulls data from the local cache memory queue according to the scheduling message according to the maximum number of tasks pulled at a time by the local cache memory queue and sends it to the worker thread pool; Re-sending the target data to the local cache memory queue, and re-determining whether the local cache memory queue is full; If the target data fails to be sent to the local cache memory queue for the first time, the target data will be stored in the MQ distributed queue, and after a preset time, the target data will be sent to the failure compensation handler, and the target data will be resubmitted and sent to the local cache memory queue through the failure compensation handler.
Citation Information
Patent Citations
Method and system for ensuring thread pool reliability under high concurrency
CN116643855A
Data persistence method and device
CN117056403A