A method and apparatus for data processing
By using coroutines in data processing for batch data updates, the problems of waste of system resources and low performance in the existing technology are solved, and an efficient and real-time data processing solution is realized, and system resource utilization and database operations are optimized.
Patent Information
- Application Number
- CN202010789270.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-08-07
AI Technical Summary
The prior art consumes a lot of system resources during data processing, resulting in low performance and difficulty in meeting the requirements of Internet services for real-time and high efficiency. Frequent switching of thread scheduling leads to resource waste and memory overflow problems that are difficult to solve.
Co-processing is used for data processing. By creating target coroutines when the total number of coroutines is lower than the preset threshold, batch data updates are performed in units of data slices, and coroutines are released after the update is completed, the database operation frequency is reduced, and the dynamic expansion and contraction strategy and user-state execution characteristics of coroutines are used to reduce system resource consumption.
It improves data processing efficiency and system performance, reduces system resource consumption, optimizes database operation frequency, avoids performance bottlenecks caused by memory overflow and frequent switching, and meets the requirements of high concurrency and real-time.
Smart Images

Figure CN114064667B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method and device for data processing. Background Art
[0002] With the development of Internet technology and the popularization of Internet applications, the requirements for the efficiency of data processing by Internet services are also constantly increasing.
[0003] In the prior art, when pulling and updating data, a multi-threaded parallel processing method is usually adopted.
[0004] However, during thread scheduling, frequent switching between the kernel mode and the user mode is required, resulting in low performance. It also consumes a large amount of system resources and time. Moreover, the operation frequency of the database is also relatively high, which will further consume a large amount of system resources.
[0005] Therefore, in data processing, a data processing technical solution that can reduce the consumed system resources, improve data processing efficiency, and system performance is needed. Summary of the Invention
[0006] Embodiments of this application provide a method and device for data processing, which are used to reduce the consumed system resources, improve data processing efficiency, and system performance when performing data processing.
[0007] On the one hand, a method for data processing is provided, including:
[0008] When it is determined that there is a target data set that needs to be updated, obtain the total number of currently running coroutines;
[0009] When the total number of coroutines is lower than a preset number threshold, create a target coroutine for the target data set, and increment the total number of coroutines by 1. The preset number threshold is determined according to the resource configuration information;
[0010] Call the target coroutine, and use data slices as the data update unit to update the target data set stored in the database;
[0011] Release the target coroutine, and decrement the total number of coroutines by 1.
[0012] On the one hand, a device for data processing is provided, including:
[0013] An acquisition unit, configured to obtain the total number of currently running coroutines when it is determined that there is a target data set that needs to be updated;
[0014] A creation unit, configured to create a target coroutine for the target data set and increment the total number of coroutines by 1 when the total number of coroutines is lower than a preset number threshold. The preset number threshold is determined according to the resource configuration information;
[0015] An update unit, configured to call a target coroutine, and use a data slice as a data update unit to update target data sets stored in a database.
[0016] A release unit, configured to release the target coroutine and decrement the total number of coroutines by 1.
[0017] On the one hand, a control device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it performs the steps of any of the above data processing methods.
[0018] On the one hand, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of any of the above data processing methods.
[0019] On the one hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of any of the above data processing.
[0020] In a data processing method and apparatus provided in an embodiment of the present application, whenever it is determined that a target data set needs to be updated and the total number of currently running coroutines is lower than a preset number threshold, a target coroutine is created and the target coroutine is called. Using a data slice as a data update unit, batch data update is performed on the target data sets stored in the database, and after the current data update is completed, the target coroutine is released. In this way, using coroutines for data processing and batch data processing on the database reduces the consumed system resources and improves the data processing efficiency and system performance.
[0021] Other features and advantages of the present application will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0023] Figure 1 It is a schematic diagram of a system architecture for a data processing method in an embodiment of the present application;
[0024] Figure 2 It is a schematic flowchart of a data processing method in an embodiment of the present application;
[0025] Figure 3 It is a schematic structural diagram of a data processing device in an embodiment of the present application;
[0026] Figure 4 It is a schematic structural diagram of a control device in an embodiment of the present application. Detailed implementation manners
[0027] In order to make the objectives, technical solutions and beneficial effects of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0028] First, some terms involved in the embodiments of the present application will be described to facilitate the understanding of those skilled in the art.
[0029] Terminal device: It can be a mobile terminal, a fixed terminal or a portable terminal, such as a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system device, a personal navigation device, a personal digital assistant, an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the terminal device can support any type of interface for users (such as wearable devices), etc.
[0030] Server: It can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.
[0031] Thread: It is a kernel object of the operating system. A thread can be used as a basic unit for independent operation and independent scheduling. The role of a thread is to make full use of the Central Processing Unit (CPU) to process tasks in parallel.
[0032] Coroutine: It is a thread simulated at the application layer. In the embodiments of this application, Go coroutines are taken as an example for illustration. A Go coroutine is a Go function or method that runs concurrently with other Go coroutines in the same address space. A running program consists of one or more Go coroutines.
[0033] Coroutine channel: It is used to send and receive shared data between multiple Go coroutines, thereby achieving data synchronization. When using it, it is necessary to know the data types sent and received in the coroutine channel. A coroutine channel can be regarded as a pipeline erected between two Go coroutines. One Go coroutine can store data in this pipeline, and another Go coroutine can retrieve data from this pipeline.
[0034] Data slice: The data structure of an array slice can be abstracted into the following three variables: a pointer to the original array, the number of elements in the array slice, and the allocated storage space of the array slice. From the perspective of underlying implementation, an array slice actually still uses an array to manage elements. Based on the array, a series of management functions are added to the array slice, which can dynamically expand the storage space at any time and can be passed around arbitrarily without causing the elements it manages to be duplicated.
[0035] Mutex lock: It is used to ensure the integrity of shared data operations. After locking the shared data, only one thread can access the shared data.
[0036] Cloud storage: It is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through functions such as cluster applications, grid technology, and distributed file systems, and collaborates through application software or application interfaces to jointly provide data storage and business access functions to the outside world.
[0037] Currently, the storage method of the storage system is as follows: Create a logical volume. When creating a logical volume, physical storage space is allocated for each logical volume. This physical storage space may be composed of a disk of a certain storage device or several storage devices. The client stores data on a certain logical volume, that is, stores the data on the file system. The file system divides the data into many parts, and each part is an object. The object contains not only data but also additional information such as data identifiers. The file system writes each object into the physical storage space of the logical volume respectively, and the file system will record the storage location information of each object. Thus, when the client requests to access data, the file system can enable the client to access the data according to the storage location information of each object.
[0038] The process of the storage system allocating physical storage space for a logical volume is specifically as follows: According to the capacity estimation of the objects stored in the logical volume (this estimation usually has a large margin relative to the capacity of the objects to be actually stored) and the group of the Redundant Array of Independent Disk (RAID), the physical storage space is pre-divided into stripes, and a logical volume can be understood as a stripe, thereby allocating physical storage space for the logical volume.
[0039] Database: Briefly, it can be regarded as an electronic filing cabinet - a place for storing electronic files, where users can perform operations such as adding, querying, updating, and deleting data in the files. The so-called "database" is a data set stored together in a certain way, shared by multiple users, with as little redundancy as possible, and independent of application programs.
[0040] Database management system: It is a computer software system designed to manage databases, generally having basic functions such as storage, interception, security guarantee, and backup. Database management systems can be classified according to the database models they support, such as relational, Extensible Markup Language, or according to the types of computers they support, such as server clusters, mobile phones; or according to the query languages they use, such as Structured Query Language (SQL), XQuery; or according to the key points of performance impulse, such as the maximum scale, the highest running speed; or other classification methods. No matter which classification method is used, some database management systems can cross categories. For example, they can support multiple query languages at the same time.
[0041] The design concept of the embodiments of the present application will be introduced below.
[0042] With the development of Internet technology and the popularization of Internet applications, the amount of data is also increasing continuously. For Internet services that need to pull a large amount of data and then update the database, the requirements for the system resources consumed by data processing and the processing efficiency are also increasing continuously.
[0043] In the traditional method, when pulling data and updating the database, threads or processes are usually used to execute tasks concurrently.
[0044] However, threads need to switch frequently between the kernel mode and the user mode, which will consume a large amount of system resources and time costs, resulting in waste of resources, low performance, and low data processing efficiency. When the amount of data pulled is huge, it is difficult to meet the requirements of Internet services with high real-time requirements. Moreover, using a stack with a fixed size that cannot be dynamically changed, it is difficult to troubleshoot problems such as memory overflow.
[0045] Obviously, the traditional technology does not provide a technical solution for data processing that can reduce the consumed system resources, improve the system performance and data processing efficiency. Therefore, a technical solution for data processing is needed to reduce the consumed system resources, improve the system performance and data processing efficiency during data processing.
[0046] A coroutine is a more lightweight entity than a thread, consuming resources at the KB level. It adopts a dynamic expansion and contraction strategy to control the stack size, achieving dynamic memory control. It is not managed by the operating system kernel but is completely controlled by the program, that is, it executes in the user state and does not need to frequently switch with the kernel state. This can greatly improve the system performance and data processing efficiency, as well as reduce the consumed resources, effectively solve the problem of untimely data pulling. Moreover, the total number of coroutines can be controlled according to the CPU resources to maximize the utilization of the CPU. It can also perform batch data updates on the database to reduce the frequency of database operations, further optimizing the system performance.
[0047] Based on the above considerations and analyses, the embodiments of the present application provide a data processing solution. In this solution, whenever it is determined that a target data set needs to be updated and the total number of currently running coroutines is lower than the preset number threshold, a target coroutine is created and the target coroutine is called. Using data slices as the data update unit, batch data updates are performed on the target data set stored in the database. After this data update is completed, the target coroutine is released. In this way, using coroutines for data processing and batch data processing on the database reduces the consumed system resources and improves the data processing efficiency and system performance.
[0048] To further illustrate the technical solution provided by the embodiments of the present application, the following will be described in detail in combination with the accompanying drawings and specific implementation manners. Although the embodiments of the present application provide method operation steps as shown in the following embodiments or drawings, based on routine or non-creative labor, more or fewer operation steps may be included in the method. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present application. During the actual processing of the method or when the device executes, it can be executed in the order shown in the embodiments or drawings or executed in parallel.
[0049] Refer to Figure 1 As shown, it is a schematic diagram of the architecture of a data processing system. It includes multiple user terminals 100, a control device 101, and a source data device 102.
[0050] Control device 101: When a new data update task needs to be executed, it pulls data from the source data device 102, creates corresponding coroutines, and calls the coroutines to execute the data update task in parallel. After the data update task is completed, it performs the coroutine release operation.
[0051] Among them, the control device 101 can be a server or a terminal device, including a database for storing data, and may also include a database management system. For example, the control device 101 is a database server.
[0052] Source data device 102: used to provide source data for the control device 101.
[0053] Optionally, the source data device 102 can store data in the form of cloud storage or in other ways, and can include only one device or a collection of multiple devices. For example, the source data devices are a video data server and a news data server.
[0054] User terminal 100: can be a terminal device, used to obtain data from the control device 101 so that the user can browse web pages, etc.
[0055] In an application scenario, the control device 101 is a database server in an airplane, and the source data device 102 is a cloud server.
[0056] Since the user cannot obtain data through the Internet after the airplane takes off, the database server pulls various data from the cloud server before the airplane takes off and stores the pulled data in the database. Each time the database server pulls data, a corresponding coroutine is created, and the coroutine is called to update the pulled data to the database in the database server.
[0057] After the airplane takes off, the user can use a local area network such as Wireless-Fidelity (WIFI) inside the airplane through the user terminal 100 to obtain data such as news, short videos, movies, and TV stored in the database server, which provides convenience for the user.
[0058] The embodiments of this application are mainly applied to updating the target data set in the database of the control device 101.
[0059] Refer to Figure 2 As shown, it is a flowchart of the implementation of a data processing method provided by this application. The specific process of this method is as follows:
[0060] Step 200: The control device creates a wait group.
[0061] Specifically, the wait group is used to control the synchronous processing of multiple tasks and can ensure the completion of multiple tasks in a concurrent environment. For a group of waiting tasks, it is not necessary to create a wait group for each task separately, and only one wait group needs to be declared.
[0062] In one implementation, sync.WaitGroup represents a wait group. Declaring a wait group can be done using var wg sync.WaitGroup.
[0063] It should be noted that the data processing method in the embodiments of this application is described by taking Go language writing as an example. In actual applications, the programming language used can also be set according to the actual application scenario, which is not limited here.
[0064] Step 201: When the control device determines that the target data set needs to be updated, obtain the total number of currently running coroutines.
[0065] Specifically, when executing step 201, any one or a combination of the following two methods can be adopted:
[0066] The first method is: when the preset duration is reached, the control device determines that the target data set needs to be updated, and obtains the total number of currently running coroutines through the wait group.
[0067] It should be noted that a counter is maintained in the wait group. The initial value of the counter is 0. After creating a coroutine, the counter is incremented by 1. After releasing a coroutine, the counter is decremented by 1. When the counter is decremented to 0, the wait group is released.
[0068] That is to say, the wait group counts the total number of currently running coroutines through the counter. Therefore, the control device can obtain the total number of currently running coroutines through the wait group.
[0069] In actual applications, the preset duration can be a fixed duration or a specified time point, which is set according to the actual application scenario and is not limited here. For example, the preset duration is 1 minute.
[0070] Among them, the target data set can be the target data that the control device pre-sets and needs to be continuously updated. The control device can obtain the target data through a specified website or URL, etc.
[0071] For example, the target data is a specified TV drama, and the control device determines 10 o'clock, the update time of the TV drama, as the time to update the local TV drama data.
[0072] Another example is that the target data is the top 10 movies with the highest ratings in a video application, and the control device determines 12 o'clock every Sunday as the time to update the local movies.
[0073] The second method is as follows: when the source data device determines that the target data has changed, it sends a data update message to the control device. When the control device determines that it has received the data update message, it determines that the target data set needs to be updated, and obtains the total number of currently running coroutines through a wait group.
[0074] In practical applications, when the target data is updated irregularly, to reduce the consumed resources, the control device does not need to periodically send data update requests to the source data device. Instead, when the source data device determines that the target data has changed, it notifies the control device that data update is required.
[0075] In this way, data can be pulled periodically, and data can also be obtained when the source data changes.
[0076] In the embodiments of the present application, the wait group is used to create and release coroutines, and determine the maximum value of the total number of coroutines according to the resource configuration information. Since there is usually only one data interface, the total number of coroutines running in parallel can be controlled by the data acquisition frequency and the resource configuration information.
[0077] Step 202: The control device acquires the target data.
[0078] Specifically, the control device sends a data acquisition request to the source data device and receives the target data returned by the source data device.
[0079] Among them, the number of target data acquired by the control device each time can be one or multiple. In practical application scenarios, a large amount of target data is usually acquired to update the database.
[0080] In one implementation, the control device sends a data acquisition request to the source data device through a data interface, and receives the target data returned by the source data device through this data interface.
[0081] Among them, the data interface is an interface for acquiring data from a third party. The data type corresponding to the target data acquired at one time can be one or multiple. Each target data can correspond to one data type or multiple data types.
[0082] For example, the target data can be news data, game data, learning data, entertainment data, or data obtained through Weishi.
[0083] In this way, data can be acquired through the data interface.
[0084] Step 203: When the total number of coroutines is lower than the preset number threshold, the control device creates target coroutines for the target data set.
[0085] Specifically, when the total number of coroutines is lower than the preset number threshold, the control device creates target coroutines for the target data set through a wait group.
[0086] In one implementation, the wait group uses wg.add to increment the total number of coroutines by 1 and wg.Done to decrement the total number of coroutines by 1.
[0087] Among them, the preset number threshold is used to control the maximum value of the total number of coroutines running in parallel. The preset number threshold is determined according to the resource configuration information. The preset number threshold can be set in advance or determined in real time.
[0088] In one implementation, the preset number threshold is determined according to the number of CPU cores. The control device pre-configures the preset number threshold according to the number of CPU cores.
[0089] In this way, the preset number value remains fixed.
[0090] In one implementation, the control device obtains the current idle system resources in real time and determines the preset number threshold according to the idle system resources.
[0091] In this way, the preset number threshold can be adjusted in real time according to the real-time operating conditions of the control device, so as to maximize the utilization of system resources.
[0092] This is because if a new data update task needs to be processed, a coroutine will be created. However, as the data update speed becomes faster and faster, the number of new data update tasks that need to be processed may reach hundreds of thousands, millions or even more. If the corresponding number of coroutines is created during operation, the CPU needs to switch between a large number of coroutines, which will consume a large amount of system resources, resulting in a decrease in system processing efficiency or even system crashes. Therefore, in the embodiments of the present application, the preset number threshold is pre-configured according to resource configuration information such as hardware processing capabilities and system performance to control the total number of running coroutines.
[0093] When it is determined that the total number of coroutines is not lower than the preset number threshold, the control device stops creating target coroutines until a coroutine is released, and then the control device creates target coroutines.
[0094] Further, before creating a coroutine, the control device creates a coroutine channel (channel). When multiple coroutines run in parallel, shared data can be sent and received through the coroutine channel to achieve data synchronization.
[0095] In one implementation, after creating a coroutine, the control device increments the total number of coroutines by 1.
[0096] Further, when creating a target coroutine, the control device can also adjust the total number of coroutines and then create the coroutine. Specifically, the following steps can be adopted:
[0097] The control device obtains the total number of coroutines, adds 1 to the total number of coroutines to obtain the adjusted total number of coroutines, and determines whether the adjusted total number of coroutines is not higher than a preset number threshold. If so, it creates target coroutines for the target data set through a wait group. Otherwise, when there is a coroutine release such that the total number of coroutines is not higher than the preset number threshold, it then creates target coroutines for the target data set through the wait group.
[0098] In this way, through the resource configuration information, coroutines are created to maximize the utilization of system resources. And after each pull of the target data, a target coroutine is created. The higher the data pull frequency, the more coroutines are created. Thus, the total number of concurrently running coroutines is controlled through the preset number threshold and the update frequency of the target data.
[0099] Step 204: The control device uses a mutex lock to lock and protect each target data.
[0100] Specifically, the control device uses a mutex lock to lock and protect the target data obtained from the source data device.
[0101] When multiple coroutines are processed in parallel, each coroutine locks before operating on the corresponding target data. Only after successfully locking can it operate, and unlocks after the operation. Through the "lock", the access to resources becomes a mutually exclusive operation, so that the locked target data will not be operated on by multiple coroutines simultaneously, and time-related errors will no longer occur, thus avoiding data chaos.
[0102] In this way, after locking and protecting the target data through the mutex lock, only the target coroutines can access the target data, improving the accuracy of the data.
[0103] Step 205: The control device calls the target coroutines to update the target data set stored in the database with data slices as the data update unit.
[0104] Specifically, when executing step 205, the following steps can be adopted:
[0105] S2051: The control device calls the target coroutines to respectively obtain the data types corresponding to each pulled target data.
[0106] Optionally, the data type can be determined according to the data format, such as audio data, video data, image data, and text data, etc. It can be determined according to the application scope of the data, such as education data, game data, news data, and training data, etc. It can also be determined according to the index stored in the database, such as mobile phone data and computer data, etc. In practical applications, the data type can be set according to the actual application scenario, and there is no limitation here.
[0107] S2052: The control device calls the target coroutine, and adds each target data to a different data slice according to the data type of each target data.
[0108] Specifically, the control device adds each target data to the corresponding data slice in accordance with the data type in the order of the target data, so that the target data of the same data type are added to the same data slice.
[0109] In one implementation, the control device performs the following steps for each target data:
[0110] Get the data type corresponding to the target data, and determine whether there is an incomplete data slice that carries the target data of the data type. If so, add the target data to the data slice; otherwise, create a new data slice and add the target data to the new data slice.
[0111] The unfilled data slice may be created by the target coroutine or a data slice created by a coroutine that has been released in the waiting group.
[0112] Among them, the Go language provides array slices with their own data structure, which can be abstracted into the following three variables: a pointer to the native array, the number of elements in the array slice, and the storage space allocated to the array slice. The storage space can be dynamically expanded at any time and can be passed around at will without causing the managed elements to be copied repeatedly.
[0113] In this way, the control device can add all the target data obtained from the source data device to the corresponding data slice.
[0114] Furthermore, the control device may also perform data update by using an array as a data update unit.
[0115] It should be noted that the length of an array cannot be modified after it is defined; an array is a value type, and a copy will be generated each time it is passed. However, the Go language provides array slices to make up for the shortcomings of arrays.
[0116] Furthermore, the control device may also perform data statistical processing on the target data corresponding to the multiple coroutines running in parallel, and add the data obtained after the statistical processing to the corresponding data slices.
[0117] In one implementation, the control device receives designated data sent by other coroutines through a target coroutine and a corresponding coroutine channel, and calls the target coroutine to perform data statistics processing on the target data and the received designated data.
[0118] For example, assume that the data obtained by the first coroutine are {a1, b1}, and the data obtained by the second coroutine are {a2, b2}. Then, the first coroutine sends a1 and b1 to the second coroutine through the corresponding coroutine channel. The second coroutine determines the sum of a1, b1, a2, and b2, and adds the sum to the data slice.
[0119] Among them, the statistical processing of the unified data can be set according to the actual application scenario. For example, it can be to calculate the sum or average value, etc., which is not limited here.
[0120] Similarly, the control device can also send the target data to other coroutines through the target coroutine and the corresponding coroutine channel, and call other coroutines to perform data statistical processing on the target data and the specified received data.
[0121] S2053: The control device updates the database according to the multiple target data included in the data slice.
[0122] Specifically, for each data slice, the control device has the following steps:
[0123] When the data volume corresponding to the target data in the data slice reaches the preset capacity threshold, update the target data in the data slice to the database.
[0124] Among them, the preset capacity threshold is the storage space allocated to the data slice. That is to say, when the data slice is full, update the target data in the data slice to the database.
[0125] In practical applications, the preset capacity threshold can be set according to the actual application scenario, which is not limited here.
[0126] In this way, when performing batch data updates according to the data slices corresponding to the target data carrying the same data type, since the data types corresponding to the target data are the same, the data update efficiency can be improved and the time cost can be reduced.
[0127] Among them, when updating the target data in the data slice to the database, for each target data in the data slice, the control device performs the following steps:
[0128] According to the original data stored in the database, judge the operation type of the target data. If the operation type indicates expired data, determine the original data corresponding to the target data in the database and delete the original data; if the operation type indicates modified data, determine the original data corresponding to the target data in the database and replace the original data with the target data; if the operation type indicates new data, add the target data to the corresponding table in the database.
[0129] In this way, batch operations of data deletion, data modification, and new data addition to the database can be achieved.
[0130] Furthermore, before adding the target data to the database, the following steps can also be executed:
[0131] The control device determines whether the length of the table reaches a preset length threshold. If so, it deletes the earliest added original data in the table to vacate the corresponding table space for the target data.
[0132] In practical applications, the preset length threshold can be set according to the actual application scenario. For example, the preset length threshold is 15 rows, and there is no limitation here.
[0133] Furthermore, before replacing the original data with the target data or adding the target data to the database, the control device can also execute the following steps:
[0134] The control device removes the characters of the first specified type in the target data, such as HyperText Transport Protocol (HTTP) and https, etc., and inserts the characters of the second specified type, such as inserting escape characters, etc.
[0135] In practical applications, both the first specified type and the second specified type can be set according to the actual application scenario, and there is no limitation here.
[0136] In this way, after format parsing and content parsing of the pulled target data through coroutines, it can be batch stored in the database.
[0137] Furthermore, after all the target data is added to the corresponding data slices, if there is no data slice with a data volume reaching the preset capacity threshold, there is no need to update the database, then step 207 is executed, that is, directly release the target coroutine and the corresponding coroutine channel, and do not release the data slices.
[0138] In the traditional technology, every time the control device obtains a target data, it will perform a database operation. The operation frequency of the database is relatively high, which will consume a large amount of system resources and the concurrent access volume of the database is relatively large. In the embodiments of the present application, only when the data slice is full, that is, when a certain amount of data is reached, the coroutine will batch execute the corresponding database operation, rather than operating the database every time there is a piece of data. This reduces the frequent operation of the database, improves the efficiency of pulling data and database update, and reduces the consumption of system resources and concurrent access volume.
[0139] Step 206: The control device releases each data slice with a data volume reaching the preset capacity threshold.
[0140] It should be noted that the storage space allocated to the data slice is the cache in the memory.
[0141] In this way, after the data in the data slice is full, the data in the data slice is updated to the database. Then, after the update of the data slice is completed, the data slice is released to prepare for the next data pull, enabling the control device to create a new data slice when pulling data next time to add new target data, recycling resources and reducing the generation of system memory fragmentation.
[0142] Step 207: The control device releases the target coroutine.
[0143] Specifically, after the target data in each data slice with the data volume reaching the preset capacity threshold is updated to the database, that is, after the control device determines that there is no data slice with the data volume reaching the preset capacity threshold, the control device releases the target coroutine.
[0144] Furthermore, before releasing the target coroutine, the corresponding coroutine channel is released.
[0145] In the embodiment of the present application, before creating the target coroutine, the coroutine channel is created first, and before releasing the target coroutine, the corresponding coroutine channel is also released first.
[0146] It should be noted that the coroutine channel and the target coroutine are created correspondingly. Data can be sent and received between coroutines through the coroutine channel to achieve data synchronization.
[0147] Furthermore, after the control device releases the target coroutine, the total number of coroutines is decreased by 1, so that a new coroutine can be created when obtaining data subsequently.
[0148] Furthermore, after all the corresponding target data is added to the corresponding data slice, there may be data slices that are not full. At this time, the control can also update the target data in the data slices that are not full to the database, so that after the target data in each data slice is updated to the database, the data slice and the target coroutine can be released.
[0149] Step 208: The control device performs an unlocking process on each target data.
[0150] In this way, after the coroutine is released, the corresponding target data can be unlocked to avoid deadlock operations.
[0151] Step 209: The control device determines whether to terminate the data update. If so, it executes Step 210; otherwise, it executes Step 201.
[0152] Specifically, the termination of data update can be determined according to the user's instruction or according to the termination trigger condition. For example, the target data in the source database stops updating, or the number of updates reaches the specified number of updates, etc.
[0153] In one implementation, when it is determined that the total number of currently running coroutines is the specified number, the control device determines to terminate the data update.
[0154] Optionally, the specified number can be 0.
[0155] Step 210: The control device releases all the created coroutines.
[0156] Specifically, after the control device releases the channels and coroutines, since there may be channels and coroutines whose release fails, the control device performs the release operation on all the channels and coroutines created through the wait group again to ensure that all the channels and coroutines are released, thereby reducing the generation of system memory fragmentation.
[0157] Furthermore, before executing step 210, the control device also determines whether there are data slices that have not been released. Since the coroutine only releases the full data slices, there may be unfilled data slices that have not been updated and released. The control device updates the target data in the unfilled data slices to the database and releases the above-mentioned unfilled data slices.
[0158] Finally, the control device ends the wait group.
[0159] Among them, the coroutine based on the Go language has the advantages of high concurrency (thousands or tens of thousands of single-machine accesses per second), short program life cycle, high I / O and low computing, and less system resource occupation. Taking advantage of the advantages of the coroutine, the coroutines are synchronized to run, and after obtaining data from the data interface, they are pulled again. It is simple and convenient to use. The default memory occupation is much less than that of JAVA and C threads, which is 2KB by default and increases dynamically with the usage size, and is automatically released after use.
[0160] In the traditional method, it usually takes 5 hours to pull a complete set of data by using threads or processes. In the embodiments of the present application, it only takes 4 minutes to pull a complete set of data by using coroutines, which greatly improves the data pulling efficiency and can meet the services with high requirements for data content timeliness. Moreover, the peak value of the CPU is fully utilized, and the CPU resources can be maximally utilized, further improving the pulling efficiency. The resources consumed by threads are at the MB level, and frequent switching between the kernel mode and the user mode is required during scheduling, which will consume a large amount of resources. However, coroutines are a more lightweight existence than threads, consuming resources at the KB level and not being managed by the kernel of the operating system but completely controlled by the program, that is, executed in the user mode without the need for frequent switching with the kernel mode, resulting in a great performance improvement and greatly reducing the consumed resources. Usually, an ordinary server can support millions of coroutines. Further, threads usually use a fixed-size stack, while coroutines can adopt a strategy of dynamic expansion and contraction, with an initial amount of 2k and a maximum expansion to 1g, having strong flexibility, occupying less resources, and not having problems such as memory overflow and segment faults that are difficult to troubleshoot like threads.
[0161] In the embodiments of the present application, by controlling the total number of coroutines in the wait group, when the total number of coroutines is not less than the preset number threshold, creating coroutines is suspended. After it is determined through the wait group that the coroutines and channels have been fully executed, the resources corresponding to the coroutines and channels are released, and then new coroutines are created. In this way, resources such as the created coroutines and channels can be monitored, the generation of system memory fragmentation can be reduced, and the resource configuration information can be maximally utilized to optimize the system performance and improve the data processing efficiency.
[0162] Based on the same inventive concept, in the embodiments of the present application, a data processing device is further provided. Since the principle of solving problems by the above-mentioned device and equipment is similar to that of a data processing method, therefore, the implementation of the above-mentioned device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0163] As Figure 3 shown, it is a schematic structural diagram of a data processing device provided by the embodiments of the present application. A data processing device includes:
[0164] An obtaining unit 301, configured to obtain the total number of currently running coroutines when it is determined that there is a target data set that needs to be updated;
[0165] A creating unit 302, configured to create target coroutines for the target data set and increment the total number of coroutines by 1 when the total number of coroutines is lower than the preset number threshold, and the preset number threshold is determined according to the resource configuration information;
[0166] An update unit 303, configured to call a target coroutine, and use data slices as data update units to update the target data set stored in the database;
[0167] A release unit 304, configured to release the target coroutine and decrement the total number of coroutines by 1.
[0168] Preferably, the update unit 303 is configured to:
[0169] Obtain the data types corresponding to each pulled target data respectively;
[0170] Add each target data to different data slices according to the data types of the target data;
[0171] Update the database according to the multiple target data included in the data slice.
[0172] Preferably, the release unit 304 is further configured to:
[0173] When it is determined that the total number of currently running coroutines is a specified number, release all the created coroutines.
[0174] Preferably, the update unit 303 is further configured to:
[0175] Use a mutex lock to lock and protect each target data;
[0176] After the data in the target data set is updated, it further includes:
[0177] Unlock each target data.
[0178] Preferably, the update unit 303 is further configured to:
[0179] Receive specified data sent by other coroutines through the target coroutine and the corresponding coroutine channel, and call the target coroutine to perform data statistical processing on the target data and the received specified data;
[0180] Wherein, the coroutine channel is used for data communication between different coroutines.
[0181] In the method and device for data processing provided by the embodiments of the present application, whenever it is determined that the target data set is updated and the total number of currently running coroutines is lower than the preset number threshold, a target coroutine is created, and the target coroutine is called to perform batch data update on the target data set stored in the database with data slices as data update units, and after this data update is completed, the target coroutine is released. In this way, using coroutines for data processing and batch data processing of the database reduces the consumed system resources and improves the data processing efficiency and system performance.
[0182] Figure 4The structural schematic diagram of a control device 4000 is shown. Refer to Figure 4 As shown, the control device 4000 includes: a processor 4010, a memory 4020, a power supply 4030, a display unit 4040, and an input unit 4050.
[0183] The processor 4010 is the control center of the control device 4000, connecting each component through various interfaces and lines, and executing various functions of the control device 4000 by running or executing software programs and / or data stored in the memory 4020, thereby monitoring the control device 4000 as a whole.
[0184] In the embodiment of the present application, when the processor 4010 calls the computer program stored in the memory 4020, it executes the data processing method provided in the embodiment shown in Figure 2 the embodiment shown in
[0185] Optionally, the processor 4010 may include one or more processing units; preferably, the processor 4010 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, applications, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 4010. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be implemented separately on independent chips.
[0186] The memory 4020 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, various applications, etc.; the data storage area may store data created according to the use of the control device 4000, etc. In addition, the memory 4020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices, etc.
[0187] The control device 4000 further includes a power supply 4030 (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 4010 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption through the power management system.
[0188] The display unit 4040 can be used to display information input by the user or information provided to the user, as well as various menus of the control device 4000, etc. In the embodiments of the present invention, it is mainly used to display the display interfaces of various applications in the control device 4000 and objects such as text and pictures displayed in the display interfaces. The display unit 4040 may include a display panel 4041. The display panel 4041 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.
[0189] The input unit 4050 can be used to receive information such as numbers or characters input by the user. The input unit 4050 may include a touch panel 4051 and other input devices 4052. Among them, the touch panel 4051, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 4051).
[0190] Specifically, the touch panel 4051 can detect the touch operation of the user, detect the signals brought by the touch operation, convert these signals into contact coordinates, send them to the processor 4010, and receive and execute the commands sent by the processor 4010. In addition, the touch panel 4051 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. The other input devices 4052 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0191] Of course, the touch panel 4051 can cover the display panel 4041. After the touch panel 4051 detects a touch operation on or near it, it is transmitted to the processor 4010 to determine the type of touch event. Subsequently, the processor 4010 provides a corresponding visual output on the display panel 4041 according to the type of touch event. Although in Figure 4 the touch panel 4051 and the display panel 4041 are implemented as two independent components to realize the input and output functions of the control device 4000, in some embodiments, the touch panel 4051 and the display panel 4041 can be integrated to realize the input and output functions of the control device 4000.
[0192] The control device 4000 may further include one or more sensors, such as a pressure sensor, a gravitational acceleration sensor, a proximity light sensor, etc. Of course, according to the needs in specific applications, the above control device 4000 may further include other components such as a camera. Since these components are not the key components used in the embodiments of the present application, therefore, in Figure 4It is not shown in [the figure] and will not be elaborated any further.
[0193] Those skilled in the art can understand that Figure 4 merely examples of control devices, which do not constitute a limitation on the control devices, may include more or fewer components than those shown in the figure, or combine certain components, or different components.
[0194] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for data processing in any of the above method embodiments is implemented.
[0195] The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the control method for data processing in any of the above method embodiments.
[0196] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a control device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for data processing, characterized in that, Including: When it is determined that there is a target data set that needs to be updated, obtain the total number of currently running coroutines; When the total number of coroutines is lower than a preset number threshold, create a target coroutine for the target data set and increment the total number of coroutines by 1. The preset number threshold is determined according to resource configuration information; the preset number threshold is used to control the maximum value of the total number of coroutines running in parallel; Call the target coroutine to update the target data set stored in the database with data slices as the data update unit; Release the target coroutine and decrement the total number of coroutines by 1.
2. The method according to claim 1, characterized in that Calling the target coroutine to update the target data set stored in the database with data slices as the data update unit includes: Obtain the data type corresponding to each pulled target data respectively; Add each target data to different data slices according to the data types of the target data; Update the database according to the multiple target data included in the data slice.
3. The method according to claim 2, wherein Further including: When it is determined that the total number of currently running coroutines is a specified number, release all created coroutines.
4. The method according to any one of claims 1 to 3, characterized in that Before calling the target coroutine to update the target data set stored in the database with data slices as the data update unit, further including: Use a mutex to lock and protect each target data; After the target data set is updated, further including: Unlock each target data.
5. The method according to any one of claims 1 to 3, characterized in that Further including: Receive specified data sent by other coroutines through the target coroutine and the corresponding coroutine channel, and call the target coroutine to perform data statistical processing on the target data and the received specified data; Wherein, the coroutine channel is used for data communication between different coroutines.
6. A data processing device, characterized in that, Including: An acquisition unit, configured to obtain the total number of currently running coroutines when it is determined that there is a target data set that needs to be updated; A creation unit, configured to create a target coroutine for the target data set and increment the total number of coroutines by 1 when the total number of coroutines is lower than a preset number threshold. The preset number threshold is determined according to resource configuration information; the preset number threshold is used to control the maximum value of the total number of coroutines running in parallel; An update unit, configured to call the target coroutine to update the target data set stored in the database with data slices as the data update unit; A release unit, configured to release the target coroutine and decrement the total number of coroutines by 1.
7. The device according to claim 6, characterized in that, The update unit is used for: Obtain the data type corresponding to each pulled target data respectively; Add each target data to different data slices according to the data types of the target data; Update the database according to the multiple target data included in the data slice.
8. The device according to claim 7, wherein The release unit is further used for: When it is determined that the total number of currently running coroutines is a specified number, release all created coroutines.
9. The device according to any one of claims 6-8, characterized in that, The update unit is further used for: Use a mutex to lock and protect each target data; After the target data set is updated, further including: Unlock each target data.
10. The device according to any one of claims 6-8, characterized in that, The update unit is further used for: Receive the specified data sent by other coroutines through the target coroutine and the corresponding coroutine channel, and call the target coroutine to perform data statistical processing on the target data and the received specified data; Among them, the coroutine channel is used for data communication between different coroutines.
Citation Information
Patent Citations
Data processing method and device, medium and electronic device
CN109977102A
Distributed storage system data migration method and system and related components
CN111078121A
Multi-disk concurrent data migration method, system and device and readable storage medium
CN111078628A