Data Processing Method and Device
By selecting the appropriate dimensions to combine data, the problems of computing resource consumption and timeliness reduction caused by data expansion are solved, and the data volume level is reduced and timeliness is improved.
Patent Information
- Application Number
- CN202210494109.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-05-05
AI Technical Summary
The problems of increased computing resource consumption and reduced timeliness due to data expansion when processing data are difficult to effectively solve in real-time computing scenarios.
By selecting one or more of the multiple dimensions as the grouping dimensions, the data to be processed is grouped and the appropriate dimension is selected when the preset conditions are met for merging, reducing data redundancy, reducing data volume level, and improving data timeliness.
It achieves a reduction in the data volume level, improves data timeliness, reduces computing resource consumption, and improves the overall performance of data processing.
Smart Images

Figure CN114840511B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video technology. More specifically, the present disclosure relates to a data processing method and apparatus. Background Art
[0002] With the rise of mobile networks, there is more and more data available for processing. However, when a set of data is subjected to different processes, the set of data will be copied into multiple sets of data respectively for different processes, which leads to an expansion of the data volume. Moreover, the time cost invested in processing these data and the computing resources required for these data to be calculated will also increase. It is necessary to ensure data timeliness without increasing computing resources to solve the problem of data expansion in large-scale real-time computing. Summary of the Invention
[0003] Exemplary embodiments of the present disclosure are directed to providing a data processing method and apparatus to at least solve the problems in data processing in related technologies, or may not solve any of the above problems.
[0004] According to an exemplary embodiment of the present disclosure, a data processing method is provided, including: obtaining data to be processed, where the data to be processed includes multiple sets of data, and each set of data in the multiple sets of data includes data in multiple dimensions; selecting at least one dimension from the multiple dimensions; and merging the data to be processed based on the at least one dimension to obtain merged data.
[0005] Optionally, the selecting at least one dimension from the multiple dimensions may include: performing grouping processing on the data to be processed with one or more dimensions in the multiple dimensions as grouping dimensions; obtaining the number of sets of the data to be processed corresponding to each grouping dimension; and for a first grouping dimension among the grouping dimensions of each grouping dimension, when the number of sets of the data to be processed corresponding to the first grouping dimension meets a preset condition, selecting a first dimension included in the first grouping dimension as the at least one dimension.
[0006] Optionally, the preset condition may include at least one of the number of groups in the grouping being the smallest, the number of groups in the grouping being within a preset range of the number of groups, and the number of groups in the grouping being equal to a preset number of groups.
[0007] Optionally, the grouping dimension may include any combination of the dimensions in the multiple dimensions.
[0008] Optionally, the data to be processed may include a first set of data and a second set of data, and the performing grouping processing on the data to be processed may include: for a second dimension included in a second grouping dimension, if a first data corresponding to the first set of data is the same as a second data corresponding to the second set of data, then dividing the first set of data and the second set of data into the same group.
[0009] Optionally, the at least one dimension may include a service dimension and / or a grouping dimension.
[0010] Optionally, the merging of the data to be processed based on the at least one dimension may include: merging multiple pieces of data in the data to be processed with the same service dimension data and / or the same grouping dimension data into one piece of data.
[0011] Optionally, the obtaining of the data to be processed may include: when it is determined that the local cache overflows, obtaining the data in the local cache as the data to be processed.
[0012] Optionally, the overflow threshold may be determined based on requirements for data timeliness and requirements for computing resource consumption.
[0013] Optionally, the overflow threshold may be determined based on a first coefficient for data timeliness requirements, a second coefficient for computing resource consumption requirements, a first relationship between the data timeliness requirements and the overflow threshold, and a second relationship between the computing resource consumption requirements and the overflow threshold.
[0014] Optionally, the data processing method may further include: when the merging of the data to be processed is interrupted, sending the data to be processed and a merging instruction to a preset storage device, where the merging instruction is used to instruct the preset storage device to select at least one dimension from the multiple dimensions and perform merging on the data to be processed based on the at least one dimension to obtain merged data.
[0015] According to an exemplary embodiment of the present disclosure, there is provided a data processing apparatus, including: a data acquisition unit configured to acquire data to be processed, where the data to be processed includes multiple groups of data, and each group of data in the multiple groups of data includes data of multiple dimensions; a dimension selection unit configured to select at least one dimension from the multiple dimensions; and a data merging unit configured to perform merging on the data to be processed based on the at least one dimension to obtain merged data.
[0016] Optionally, the dimension selection unit may be configured to: perform grouping processing on the data to be processed using one or more of the multiple dimensions as grouping dimensions; obtain the number of groups of the data to be processed corresponding to each grouping dimension; and for a first grouping dimension among the respective grouping dimensions, when the number of groups of the data to be processed corresponding to the first grouping dimension meets a preset condition, select a first dimension included in the first grouping dimension as the at least one dimension.
[0017] Optionally, the dimension selection unit may be configured to: the preset condition may include at least one of the number of groups in the grouping being the smallest, the number of groups in the grouping being within a preset range of the number of groups, and the number of groups in the grouping being equal to a preset number of groups.
[0018] Optionally, the grouping dimension may include a combination of any of the plurality of dimensions.
[0019] Optionally, the data to be processed may include a first set of data and a second set of data, and the dimension selection unit may be configured to: for a second dimension included in a second grouping dimension, if a first data corresponding to the first set of data is the same as a second data corresponding to the second set of data, then divide the first set of data and the second set of data into the same group.
[0020] Optionally, the at least one dimension includes a service dimension and / or a grouping dimension.
[0021] Optionally, the data merging unit may be configured to: merge multiple pieces of data in the data to be processed with the same service dimension data and / or the same grouping dimension data into one piece of data.
[0022] Optionally, the data acquisition unit may be configured to: when it is determined that the local cache overflows, acquire the data in the local cache as the data to be processed.
[0023] Optionally, the overflow threshold may be determined based on data timeliness requirements and computing resource consumption requirements.
[0024] Optionally, the overflow threshold may be determined based on a first coefficient of data timeliness requirements, a second coefficient of computing resource consumption requirements, a first relationship between data timeliness requirements and the overflow threshold, and a second relationship between computing resource consumption requirements and the overflow threshold.
[0025] Optionally, the data processing device may further include: a sending unit configured to, when the merging of the data to be processed is interrupted, send the data to be processed and a merging instruction to a preset storage device, where the merging instruction is used to instruct the preset storage device to select at least one dimension from the plurality of dimensions and merge the data to be processed based on the at least one dimension to obtain merged data.
[0026] According to an exemplary embodiment of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the data processing method according to the exemplary embodiment of the present disclosure.
[0027] According to an exemplary embodiment of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor of an electronic device, the electronic device is caused to execute the data processing method according to the exemplary embodiment of the present disclosure.
[0028] According to an exemplary embodiment of the present disclosure, there is provided a computer program product including computer programs / instructions which, when executed by a processor, implement a data processing method according to an exemplary embodiment of the present disclosure.
[0029] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0030] It is possible to achieve a reduction in the data volume level, thereby improving data timeliness and reducing the consumption of computing resources for data.
[0031] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0033] Figure 1 An exemplary system architecture in which the exemplary embodiments of the present disclosure can be applied is shown.
[0034] Figure 2 A flowchart showing a data processing method according to an exemplary embodiment of the present disclosure is shown.
[0035] Figure 3 A block diagram showing a data processing apparatus according to an exemplary embodiment of the present disclosure is shown.
[0036] Figure 4 is a block diagram of an electronic device 400 according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0037] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings.
[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0039] It should be noted here that "at least one of several items" as used in this disclosure means that it includes three parallel cases: "any one of the several items", "any combination of several items among them", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing step one and step two.
[0040] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data that have been authorized by the user or fully authorized by all parties.
[0041] In the related art, since the original data is segmented at the day level or the hour level, this results in the timeliness of data processing not being very timely, which determines that it cannot be used in some real-time interaction scenarios. Because the same data can be divided into multiple services, multiple groups, and a reference group, this causes a piece of data to be infinitely magnified, and the magnification factor depends on the number of services and the number of groups. Therefore, when actually processing data, the amount level of the data will increase by many times. Due to the increase in the amount of data, the delay in data processing is further increased, that is, the timeliness of data processing in the related art is not high, and at least the delay is at the hour level. If real-time computing is used to calculate in order to improve the overall timeliness of data processing, due to the data expansion, the real-time resources consumed will also increase exponentially.
[0042] Next, Figures 1 to 4 Specifically describe the data processing method and apparatus according to the exemplary embodiments of the present disclosure.
[0043] Figure 1 Show the exemplary system architecture 100 to which the exemplary embodiments of the present disclosure can be applied.
[0044] As Figure 1As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. Users can use the terminal devices 101, 102, 103 to interact with the server 105 via the network 104 to receive or send messages (e.g., data processing requests), etc. Various applications may be installed on the terminal devices 101, 102, 103. The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are software, they can be installed in the aforementioned electronic devices, which can be implemented as multiple software or software modules (e.g., for providing distributed services), or can be implemented as a single software or software module. Specific limitations are not imposed herein.
[0045] The server 105 may be a server that provides various services, such as a backend server that supports the applications installed on the terminal devices 101, 102, 103. The backend server can parse, store, and process data such as the received data processing requests, and feedback the data processing results to the terminal devices 101, 102, 103.
[0046] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (e.g., for providing distributed services), or can be implemented as a single software or software module. Specific limitations are not imposed herein.
[0047] It should be noted that the data processing method provided by the embodiments of the present disclosure is usually executed by the terminal device, but can also be executed by the server, or can be executed by the cooperation of the terminal device and the server. Correspondingly, the data processing device can be set in the terminal device, the server, or set in both the terminal device and the server.
[0048] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0049] When different processes are performed on input data in real time, a lot of redundant data will be generated after the input data has been processed differently. This part of the redundant data will finally be aggregated according to the same service (also known as the world in this field) and the same grouping. Data shuffle will occur during the aggregation. Since data shuffle is a resource-consuming processing flow, the present disclosure mainly solves the problem that too much data for data shuffle leads to a decrease in the overall processing performance without adding any computing resources. The specific solution is to perform an in-core aggregation locally and then perform data shuffle. In this way, the amount of data for data shuffle is reduced by at least one order of magnitude, thereby reducing the computing resources required for the data.
[0050] Figure 2 FIG. shows a flowchart of a data processing method according to an exemplary embodiment of the present disclosure. Figure 2 The data processing method in can be applied to a device running a computing engine, and particularly applicable to a device running a real-time computing engine. The real-time computing engine is, for example, but not limited to, a Flink computing engine, a Spark computing engine, etc.
[0051] Refer to Figure 2 In step S201, the data to be processed is obtained. Here, the data to be processed includes multiple groups of data, and each group of data in the multiple groups of data includes data of multiple dimensions. In one embodiment, the dimensions of different groups have the same number of levels.
[0052] In an exemplary embodiment of the present disclosure, when obtaining the data to be processed, the data in the local cache can be obtained as the data to be processed in the case where it is determined that the local cache overflows, so that data processing can be performed when the local cache overflows. In one example, data processing may not be performed when the local cache does not overflow. Here, the data to be processed can be put into the local cache. In addition, before putting the data to be processed into the local cache, the local cache can be initialized.
[0053] In an exemplary embodiment of the present disclosure, the overflow threshold is determined based on the data timeliness requirement and the computing resource consumption requirement, so as to improve the accuracy of the overflow threshold. In one embodiment, when setting the overflow threshold based on the data timeliness requirement and the computing resource consumption requirement, the overflow threshold can be determined based on a first coefficient of the data timeliness requirement, a second coefficient of the computing resource consumption requirement, a first relationship between the data timeliness requirement and the overflow threshold, and a second relationship between the computing resource consumption requirement and the overflow threshold, so as to adjust the data timeliness and the computing resource consumption requirement by setting an appropriate overflow threshold. Here, the data timeliness requirement may be inversely related to the overflow threshold, and the computing resource consumption requirement may be positively related to the overflow threshold. For example, where K is a preset value.
[0054] In an exemplary embodiment of the present disclosure, the data to be processed is obtained by performing multiple different processes on the input data respectively. For example, an AB test can be performed on the input data, so as to facilitate performing 2 processes on the input data, that is, executing 2 services.
[0055] In an exemplary embodiment of the present disclosure, the at least one dimension may include a service dimension and / or a grouping dimension, so as to merge the data to be processed based on the service dimension and / or the grouping dimension to improve the merging effect. The dimension of the data to be processed may be related to the content of the data. In one example, the dimension of the data to be processed may include, for example, a user ID.
[0056] In step S202, at least one dimension is selected from the multiple dimensions.
[0057] In an exemplary embodiment of the present disclosure, when selecting at least one dimension from the multiple dimensions, one or more dimensions in the multiple dimensions may be first used as grouping dimensions to perform grouping processing on the data to be processed, and the number of groups of the data to be processed corresponding to each grouping dimension is obtained. Then, for the first grouping dimension of each grouping dimension, when the number of groups of the data to be processed corresponding to the first grouping dimension meets a preset condition, the first dimension included in the first grouping dimension is selected as the at least one dimension, so as to improve the accuracy of the dimension for merging, and further improve the data processing effect.
[0058] In an exemplary embodiment of the present disclosure, the preset condition may include at least one of the number of groups in the grouping being the smallest, the number of groups in the grouping being within a preset range of the number of groups, and the number of groups in the grouping being equal to the preset number of groups, so as to improve the accuracy of the dimension for merging.
[0059] In an exemplary embodiment of the present disclosure, the grouping dimension may include a combination of any of the multiple dimensions, so as to perform grouping according to the required situation. For example, when the multiple dimensions include Dimension A, Dimension B, Dimension C, and Dimension D, the grouping dimension may be Dimension A (or Dimension B, or Dimension C, or Dimension D), or it may be Dimension A and Dimension B (or Dimension A and Dimension C, or Dimension A and Dimension D, or Dimension B and Dimension C, or Dimension B and Dimension D, or Dimension C and Dimension D). In this case, if the grouping dimension is Dimension A, the data in the to-be-processed data with the same Dimension A data are grouped together. For example, the data with Dimension A being world1 are grouped into one group, and the data with Dimension A being world2 are grouped into one group. If the grouping dimension is Dimension A and Dimension B, the data in the to-be-processed data with the same Dimension A data and the same Dimension B data are grouped together. For example, the data with Dimension A being world1 and Dimension B being group1 are grouped into one group, and the data with Dimension A being world2 and Dimension B being group2 are grouped into one group.
[0060] In an exemplary embodiment of the present disclosure, when grouping the multiple sets of data, the data sets with the same data in the one or more dimensions may be grouped into the same group. If the to-be-processed data includes a first set of data and a second set of data, when performing the grouping process on the to-be-processed data, for the second dimension included in the second grouping dimension, if the corresponding first data in the first set of data is the same as the corresponding second data in the second set of data, the first set of data and the second set of data are divided into the same group, so that the data in the same group have the same second data.
[0061] In step S203, the to-be-processed data is merged based on the at least one dimension to obtain the merged data. Here, the merged data can be used to replace the to-be-processed data for subsequent processing.
[0062] In an exemplary embodiment of the present disclosure, when merging the to-be-processed data based on the at least one dimension, multiple data with the same business dimension data and / or the same grouping dimension data in the to-be-processed data may be merged into one piece of data, so as to combine like terms for the to-be-processed data. Here, when merging, the data with the same dimension can be merged. In addition, other data merging methods can also be used.
[0063] In this case, if the grouping dimension is Dimension A, the data in the data to be processed with the same Dimension A data are merged into one piece of data. For example, the data with Dimension A being world1 are merged into one piece of data, and the data with Dimension A being world2 are merged into one piece of data. If the grouping dimensions are Dimension A and Dimension B, the data in the data to be processed with the same Dimension A data and the same Dimension B data are merged into one piece of data. For example, the data with Dimension A being world1 and Dimension B being group1 are merged into one piece of data, and the data with Dimension A being world2 and Dimension B being group2 are merged into one piece of data.
[0064] As an example, the following data to be processed includes 8 groups of data with 4 dimensions, and the 4 dimensions are the uid dimension, the play_duration dimension, the world dimension, and the group dimension.
[0065] (uid1,play_duration1,world1,group1)
[0066] (uid1,play_duration1,world2,group2)
[0067] (uid2,play_duration2,world1,group1)
[0068] (uid2,play_duration2,world2,group2)
[0069] (uid3,play_duration3,world1,group1)
[0070] (uid3,play_duration3,world2,group2)
[0071] (uid4,play_duration4,world1,group1)
[0072] (uid4,play_duration4,world2,group1)
[0073] If the data to be processed is merged based on the world dimension and the group dimension, the merged data is as follows:
[0074] (world1,group1, sum(play_duration1,2,3,4))
[0075] (world2, group2, sum(play_duration1, 2, 3, 4))
[0076] It can be seen that the amount of data after merging is reduced to one - fourth of the amount of data before merging.
[0077] In addition, after step S202, data flushing can be performed on the processed data. Here, since the amount of data after merging is reduced, the computing resources required for data flushing of the merged data are also correspondingly reduced.
[0078] In an exemplary embodiment of the present disclosure, when the merging of the to - be - processed data is interrupted, the to - be - processed data and a merging instruction may be sent to a preset storage device. The merging instruction is used to instruct the preset storage device to select at least one dimension from the multiple dimensions and merge the to - be - processed data based on the at least one dimension to obtain the merged data. That is, the to - be - processed data is stored in a memory (i.e., the to - be - processed data is persisted), so as to re - merge the to - be - processed data stored in the memory, thereby avoiding the loss of cached data. For example, the to - be - processed data can be persisted to a memory (such as a disk) to supplement missing data based on the to - be - processed data persisted to the memory (such as a disk) when the data merging is accidentally interrupted.
[0079] According to the data processing method of the exemplary embodiment of the present disclosure, the data timeliness can be at a delay level of minutes, such as delays of 1 minute, 2 minutes, 3 minutes, 4 minutes, 5 minutes, etc. Compared with the delay level of days or hours in the related art, the timeliness of data processing is improved.
[0080] The above has been combined with Figure 2 to describe the data processing method according to the exemplary embodiment of the present disclosure. In the following, reference will be made to Figure 3 to describe the data processing device and its units according to the exemplary embodiment of the present disclosure.
[0081] Figure 3 A block diagram showing a data processing device according to an exemplary embodiment of the present disclosure.
[0082] Referring to Figure 3 , the data processing device includes a data acquisition unit 31, a dimension selection unit 32, and a data merging unit 33.
[0083] The data acquisition unit 31 is configured to acquire to - be - processed data. Here, the to - be - processed data includes multiple groups of data, and each group of data in the multiple groups of data includes data of multiple dimensions. Each group of data in the multiple groups of data has the same dimensions.
[0084] In an exemplary embodiment of the present disclosure, the at least one dimension includes a service dimension and / or a grouping dimension.
[0085] In an exemplary embodiment of the present disclosure, the data acquisition unit 31 may be configured to: when it is determined that the local cache overflows, acquire the data in the local cache as the data to be processed.
[0086] In an exemplary embodiment of the present disclosure, the overflow threshold may be determined based on the requirements for data timeliness and the consumption of computing resources.
[0087] In an exemplary embodiment of the present disclosure, the overflow threshold is determined based on a first coefficient of the data timeliness requirement, a second coefficient of the computing resource consumption requirement, a first relationship between the data timeliness requirement and the overflow threshold, and a second relationship between the computing resource consumption requirement and the overflow threshold. Here, the data timeliness requirement is inversely related to the overflow threshold, and the computing resource consumption requirement is positively related to the overflow threshold.
[0088] The dimension selection unit 32 is configured to select at least one dimension from the multiple dimensions.
[0089] In an exemplary embodiment of the present disclosure, the dimension selection unit 32 may be configured to: use one or more of the multiple dimensions as grouping dimensions to perform grouping processing on the data to be processed; acquire the number of groups of the data to be processed corresponding to each grouping dimension; for a first grouping dimension among the respective grouping dimensions, when the number of groups of the data to be processed corresponding to the first grouping dimension meets a preset condition, select a first dimension included in the first grouping dimension as the at least one dimension.
[0090] In an exemplary embodiment of the present disclosure, the preset condition may include at least one of the number of groups in the grouping being the smallest, the number of groups in the grouping being within a preset range of the number of groups, and the number of groups in the grouping being equal to a preset number of groups.
[0091] In an exemplary embodiment of the present disclosure, the grouping dimension includes any combination of the multiple dimensions.
[0092] In an exemplary embodiment of the present disclosure, the data to be processed may include a first group of data and a second group of data, and the dimension selection unit 32 may be configured to: for a second dimension included in a second grouping dimension, if a first data corresponding to the first group of data is the same as a second data corresponding to the second group of data, divide the first group of data and the second group of data into the same group.
[0093] The data merging unit 33 is configured to merge the data to be processed based on the at least one dimension to obtain the merged data. Here, the merged data can be used to replace the data to be processed for subsequent processing.
[0094] In an exemplary embodiment of the present disclosure, the data merging unit 33 may be configured to: merge multiple pieces of data in the data to be processed that have the same business dimension data and / or the same grouping dimension data into one piece of data.
[0095] In addition, the merged data may be subjected to data scrubbing.
[0096] In an exemplary embodiment of the present disclosure, the data processing apparatus may further include: a sending unit (not shown), configured to, when the merging of the data to be processed is interrupted, send the data to be processed and a merging instruction to a preset storage device, where the merging instruction is used to instruct the preset storage device to select at least one dimension from the multiple dimensions and merge the data to be processed based on the at least one dimension to obtain the merged data.
[0097] Regarding the apparatus in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0098] The above has been combined with Figure 3 to describe the data processing apparatus according to the exemplary embodiments of the present disclosure. Next, in combination with Figure 4 to describe the electronic device according to the exemplary embodiments of the present disclosure.
[0099] Figure 4 is a block diagram of an electronic device 400 according to the exemplary embodiments of the present disclosure.
[0100] Referring to Figure 4 , the electronic device 400 includes at least one memory 401 and at least one processor 402. A set of computer-executable instructions is stored in the at least one memory 401. When the set of computer-executable instructions is executed by the at least one processor 402, a method for data processing according to the exemplary embodiments of the present disclosure is executed.
[0101] In an exemplary embodiment of the present disclosure, the electronic device 400 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 400 does not have to be a single electronic device, and may also be any aggregate of devices or circuits capable of executing the above instructions (or instruction sets) alone or jointly. The electronic device 400 may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device that is interconnected with a local or remote (e.g., via wireless transmission) interface.
[0102] In the electronic device 400, the processor 402 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.
[0103] The processor 402 may execute instructions or code stored in the memory 401, where the memory 401 may also store data. The instructions and data may also be sent and received via a network interface device over a network, where the network interface device may employ any known transmission protocol.
[0104] The memory 401 may be integrated with the processor 402. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. In addition, the memory 401 may include a separate device, such as an external disk drive, a storage array, or other storage devices usable by any database system. The memory 401 and the processor 402 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 402 can read files stored in the memory.
[0105] In addition, the electronic device 400 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 400 may be connected to each other via a bus and / or a network.
[0106] According to an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium including instructions, such as the memory 401 including instructions, and the above instructions may be executed by the processor 402 of the device 400 to complete the above method. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0107] According to an exemplary embodiment of the present disclosure, there may also be provided a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, a method for data processing according to an exemplary embodiment of the present disclosure is implemented.
[0108] The above has been described with reference to Figures 1 to 4 a data processing method and apparatus according to an exemplary embodiment of the present disclosure. However, it should be understood that: Figure 3 the data processing apparatus and its units shown therein may be respectively configured as software, hardware, firmware, or any combination of the above items for performing specific functions, Figure 4The electronic device shown is not limited to including the components shown above, but some components can be added or deleted as needed, and the above components can also be combined.
[0109] According to the data processing method and device of the present disclosure, first, data to be processed is obtained, wherein the data to be processed includes multiple groups of data, and each group of data in the multiple groups of data includes data in multiple dimensions. Then, at least one dimension is selected from the multiple dimensions. After that, the data to be processed is merged based on the at least one dimension to obtain the merged data, so as to reduce the data volume level, improve the data timeliness, and reduce the consumption of computing resources for the data.
[0110] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0111] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A data processing method, characterized in that, Including: Obtain data to be processed, where the data to be processed includes multiple groups of data, and each group of data in the multiple groups of data includes data of multiple dimensions; Using one or more of the multiple dimensions as grouping dimensions, group the components with the same data in the one or more dimensions into the same group; Obtain the number of groups of the data to be processed corresponding to each grouping dimension; For the first grouping dimension among the grouping dimensions, when the number of groups of the data to be processed corresponding to the first grouping dimension meets a preset condition, select the first dimension included in the first grouping dimension as at least one dimension; Based on the at least one dimension, merge the data with the same dimension in the data to be processed to obtain the merged data, where, when the data to be processed includes the first group of data and the second group of data, the grouping of the components with the same data in the one or more dimensions into the same group includes: For the second dimension included in the second grouping dimension, if the first data corresponding in the first group of data is the same as the second data corresponding in the second group of data, then divide the first group of data and the second group of data into the same group.
2. The data processing method according to claim 1, wherein The preset condition includes at least one of the minimum number of groups in the grouping, the number of groups in the grouping being within a preset range of the number of groups, and the number of groups in the grouping being equal to the preset number of groups.
3. The data processing method according to claim 2, wherein The grouping dimension includes any combination of the multiple dimensions.
4. The data processing method according to claim 1, wherein The at least one dimension includes a business dimension and / or a grouping dimension.
5. The data processing method according to claim 4, wherein The merging of the data with the same dimension in the data to be processed based on the at least one dimension includes: Merging multiple data items with the same dimension in the groups with the same business dimension data and / or the same grouping dimension data in the data to be processed into one data item.
6. The data processing method according to claim 1, wherein The obtaining of the data to be processed includes: When it is determined that the local cache overflows, obtain the data in the local cache as the data to be processed.
7. The data processing method according to claim 6, characterized in that, The overflow threshold is determined based on the requirements of data timeliness and computing resource consumption.
8. The data processing method according to claim 7, wherein The overflow threshold is determined based on a first coefficient of the data timeliness requirement, a second coefficient of the computing resource consumption requirement, a first relationship between the data timeliness requirement and the overflow threshold, and a second relationship between the computing resource consumption requirement and the overflow threshold.
9. The data processing method according to claim 1, wherein Also including: When the merging of the data to be processed is interrupted, send the data to be processed and a merging instruction to a preset storage device, and the merging instruction is used to instruct the preset storage device to select at least one dimension from the multiple dimensions and merge the data to be processed based on the at least one dimension to obtain the merged data.
10. A data processing device, characterized in that, Including: A data acquisition unit configured to obtain data to be processed, where the data to be processed includes multiple groups of data, and each group of data in the multiple groups of data includes data of multiple dimensions; A dimension selection unit, configured to use one or more of the multiple dimensions as grouping dimensions, group components with the same data in the one or more dimensions into the same group, obtain the number of groups of the to-be-processed data corresponding to each grouping dimension, and for a first grouping dimension among the grouping dimensions, when the number of groups of the to-be-processed data corresponding to the first grouping dimension meets a preset condition, select a first dimension included in the first grouping dimension as at least one dimension; A data merging unit, configured to merge data with the same dimension in the to-be-processed data based on the at least one dimension to obtain merged data, wherein, when the to-be-processed data includes a first group of data and a second group of data, the dimension selection unit is configured to: For a second dimension included in a second grouping dimension, if a first data corresponding to the first group of data is the same as a second data corresponding to the second group of data, then divide the first group of data and the second group of data into the same group.
11. The data processing device according to claim 10, wherein The preset condition includes at least one of the number of groups being the smallest, the number of groups being within a preset range of the number of groups, and the number of groups being equal to a preset number of groups.
12. The data processing device according to claim 11, wherein The grouping dimension includes any combination of the multiple dimensions.
13. The data processing device according to claim 10, wherein The at least one dimension includes a service dimension and / or a grouping dimension.
14. The data processing device according to claim 13, wherein The data merging unit is configured to: Merge multiple data items with the same dimension in groups with the same service dimension data and / or the same grouping dimension data in the to-be-processed data into one data item.
15. The data processing device according to claim 10, wherein A data acquisition unit is configured to: When it is determined that the local cache overflows, acquire the data in the local cache as the to-be-processed data.
16. The data processing device according to claim 15, characterized in that, The overflow threshold is determined based on data timeliness requirements and computing resource consumption requirements.
17. The data processing device according to claim 16, characterized in that, The overflow threshold is determined based on a first coefficient of data timeliness requirements, a second coefficient of computing resource consumption requirements, a first relationship between data timeliness requirements and the overflow threshold, and a second relationship between computing resource consumption requirements and the overflow threshold.
18. The data processing device according to claim 10, characterized in that, It further includes: A sending unit, configured to send the to-be-processed data and a merge instruction to a preset storage device when the merging of the to-be-processed data is interrupted, where the merge instruction is used to instruct the preset storage device to select at least one dimension from the multiple dimensions and merge the to-be-processed data based on the at least one dimension to obtain merged data.
19. An electronic device, characterized in that, It includes: A processor; A memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the data processing method according to any one of claims 1 to 9.
20. A computer-readable storage medium stores a computer program, characterized in that, When the computer program is executed by a processor of an electronic device, the electronic device is caused to execute the data processing method according to any one of claims 1 to 9.
21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Data processing method and device, equipment and storage medium
CN113761018A