Data processing method and device, electronic equipment and storage medium

By determining the write position of state change data based on the time difference value in the data warehouse table and converting it into binary form to fill the data, the complex and inflexible problems of traditional methods are solved, efficient data management and query are achieved, and the accuracy and speed of data processing are improved.

CN120067115APending Publication Date: 2025-05-30BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510059077.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional dimensional change table design method is complex and difficult to operate, lacks flexibility, and is difficult to adapt to rapidly changing business needs, resulting in difficulty in updating and maintaining data warehouses, affecting the speed and accuracy of data processing.

Method used

By obtaining the slowly changing target data in the data warehouse table and its starting time of state change, determining the starting point and end point position of the state change data written based on the difference between the set time and the start time, the state change data is converted into binary form and filled in the corresponding position, and the data warehouse table is updated.

Benefits of technology

It realizes efficient data management and query, reduces the demand for data storage space, improves the speed and accuracy of data updates, simplifies the data management process, and provides a more reliable foundation for data analysis and decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067115A_ABST
    Figure CN120067115A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, in particular to the technical field of data warehouse technology, big data processing and the like. According to the specific implementation scheme, a data warehouse table is obtained, wherein the data warehouse table comprises target data which changes slowly along with the time state; acquiring starting time when the state of the target data changes; in response to a state change data writing request for the target data, determining a starting point position and an ending point position of state change data writing according to a difference value between the set time and the starting time; and if the sequence length corresponding to the starting point position is smaller than or equal to the current sequence length of the target data, converting the state change data into a binary form, and filling a conversion result into an interval range from the starting point position to the end point position to obtain an updated data warehouse table. According to the method and the device, the requirement of a data storage space is remarkably reduced, and the data updating speed and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, specifically to the technical fields of data warehouse technology and big data processing, and particularly to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] In the fields of data management and big data analysis, the construction and maintenance of data warehouses are one of the core tasks. However, with the continuous growth and complexity of business requirements, the types and change patterns of data have become increasingly diverse. Especially for Slowly Changing Dimension (SCD) data, this type of data requires special processing due to the changes in its status over different time periods. In order to accurately record and analyze these changes, it is usually necessary to design dedicated data warehouse tables for them.

[0003] However, the traditional dimension change table design methods face many challenges. First, these methods are often complex and difficult to operate, resulting in low efficiency for data developers in the design and implementation process. Second, due to the lack of sufficient flexibility, traditional methods are difficult to adapt to rapidly changing business requirements, making the update and maintenance of data warehouses difficult. These problems not only affect the speed and accuracy of data processing but also limit the potential of data warehouses in supporting decision-making analysis. Therefore, developing a new method that can efficiently manage and query slowly changing dimension data is of great significance for improving data processing capabilities and supporting business decisions. Summary of the Invention

[0004] The present disclosure provides a data processing method, apparatus, electronic device, and storage medium.

[0005] According to one aspect of the present disclosure, a data processing method is provided, the method comprising:

[0006] Obtaining a data warehouse table, where the data warehouse table contains target data whose status changes slowly over time;

[0007] Obtaining the starting time when the status of the target data changes;

[0008] In response to a status change data writing request for the target data, determining a starting position and an ending position for writing the status change data according to the difference between a set time and the starting time;

[0009] If the sequence length corresponding to the starting position is less than or equal to the current sequence length of the target data, converting the status change data into a binary form and filling the conversion result into the interval range between the starting position and the ending position to obtain an updated data warehouse table.

[0010] According to another aspect of the present disclosure, there is provided a data processing apparatus, the apparatus comprising:

[0011] A first acquisition module, configured to acquire a data warehouse table, where the data warehouse table contains target data whose state changes slowly over time;

[0012] A second acquisition module, configured to acquire the starting time when the state of the target data changes;

[0013] A response module, configured to, in response to a request for writing state change data for the target data, determine a starting position and an ending position for writing the state change data according to the difference between a set time and the starting time;

[0014] An update module, configured to, if the sequence length corresponding to the starting position is less than or equal to the current sequence length of the target data, convert the state change data into a binary form, and fill the conversion result into the interval range between the starting position and the ending position to obtain an updated data warehouse table.

[0015] According to a third aspect of the present disclosure, there is provided an electronic device, comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of the above technical solutions.

[0019] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in any one of the above technical solutions.

[0020] According to a fifth aspect of the present disclosure, there is provided a computer program product, comprising a computer program, where the computer program implements the method described in any one of the above technical solutions when executed by a processor.

[0021] The present disclosure provides a data processing method, apparatus, device, and storage medium. By creating a data warehouse table containing target data with slow state changes over time, the present disclosure achieves efficient data management. When a state change data write request for the target data is received, the start and end positions of data writing are accurately determined based on the difference between the set time and the start time. Subsequently, the state change data is efficiently converted into binary form and filled into the corresponding positions. This binary representation is more compact than traditional storage methods, significantly reducing the requirement for data storage space. In addition, by directly locating and filling data, unnecessary data processing steps are avoided, thereby improving the speed and accuracy of data updates. This method not only simplifies the data management process but also provides a more reliable basis for data analysis and decision support, ensuring the efficiency and accuracy of data processing.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0024] Figure 1 is a schematic diagram of the steps of the data processing method in the embodiments of the present disclosure;

[0025] Figure 2 is a schematic diagram of the process of writing data into the data warehouse table in the embodiments of the present disclosure;

[0026] Figure 3 is a schematic diagram of inserting a coding segment into a bit sequence in the embodiments of the present disclosure;

[0027] Figure 4 is a schematic diagram of the process of querying data in the data warehouse table in the embodiments of the present disclosure;

[0028] Figure 5 is a schematic diagram of obtaining a decoded segment from a bit sequence in the embodiments of the present disclosure;

[0029] Figure 6 Schematic block diagram of the data processing apparatus in the embodiments of the present disclosure;

[0030] Figure 7 is a block diagram of an electronic device for implementing the data processing method in the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0032] In the prior art, there are mainly two solutions for processing slowly changing dimension data: full snapshot tables and zipper tables. Specifically, full snapshot tables store data by making a complete backup of the entire data set at a specific point in time. This method is usually performed regularly and is suitable for scenarios where full historical data needs to be retained. However, this method has disadvantages such as data redundancy, complex queries, and performance issues. Since each snapshot is full data, it results in a large storage space requirement, and when querying historical data, it is necessary to compare and calculate among multiple table partitions, resulting in low query efficiency. In addition, regular full extraction and storage may cause greater pressure on system performance, especially in the case of large amounts of data.

[0033] Zipper tables record each change of data by adding new rows. Each row of data includes a valid start time and an end time, and is suitable for scenarios where it is necessary to track the historical changes of data and reduce storage space. However, zipper tables also have disadvantages such as complex maintenance, performance issues, and data redundancy. This method requires maintaining the valid time interval of each row of data, the data update logic is complex, and over time, the amount of data in the table will continue to increase, affecting query performance. Although it saves more space than full snapshot tables, there will still be a certain degree of data redundancy because each change adds a new row of data.

[0034] In summary, although the prior art provides some solutions for processing slowly changing dimension data, there are still problems such as low storage efficiency, complex queries, difficult maintenance, and performance bottlenecks, which limit the efficiency and flexibility of data processing.

[0035] To solve the above technical problems, the present disclosure provides a data processing method, as shown in Figure 1 shown Figure 1 is a schematic diagram of the steps of the data processing method in the embodiments of the present disclosure. This method is applied to the server side and includes:

[0036] Step S101, obtain a data warehouse table, where the data warehouse table contains target data whose status changes slowly over time.

[0037] Specifically, a data warehouse table is the basic storage unit in a data warehouse, which is used to organize and store data that has been cleaned, transformed, and aggregated for easy analysis and reporting. In this solution, the target data contained in the data warehouse table is data with slow-changing time states. This type of data usually refers to data with a low change frequency within a certain time range, such as the inventory status of goods, the activity of users, etc. The characteristic of this data is that it remains stable for a long time, but may change its state at certain specific time points. To effectively manage and query this slowly changing data, this solution records the historical state changes of each data point in the data warehouse table, including the time point of the change and the state value after the change, so as to achieve precise tracking and historical query of data changes.

[0038] Step S102: Obtain the starting time when the status of the target data changes.

[0039] Specifically, obtaining the starting time when the status of the target data changes means determining the specific time point when the data changes from one state to another. In data management and analysis, the starting time is a key timestamp used to mark the start of a data status change. This solution can accurately capture the starting time of the target data status change by monitoring and recording data changes. The specific implementation methods of this solution include: continuously monitoring the data source, and when it detects that the data status changes, immediately record the time point of this change as the starting time. This process usually involves real-time processing of data streams and event detection mechanisms to ensure that changes can be captured and recorded at the first moment of data change. In this way, by obtaining the starting time, it is beneficial to further analyze the data change trend, duration, and correlation with other data, providing an important time reference basis for subsequent data analysis, report generation, and decision support, thereby improving the accuracy and timeliness of data processing.

[0040] Step S103: In response to a request to write status change data for the target data, determine the starting position and ending position for writing the status change data according to the difference between the set time and the starting time.

[0041] Specifically, a status change data write request for target data means that when an instruction is received to record the status change of target data at a specific time point into a data warehouse table, a series of operations will be performed to process this request. Specifically, first, a time difference is calculated based on the time specified in the request and the start time of the previously recorded status change of the target data. This time difference is used to determine the start position and end position in the data warehouse table where the status change data should be written. The start position refers to the position where the status change data starts to be inserted in the bit sequence, and the end position is the position where the data insertion ends. In this way, the status change information of the target data at the set time can be accurately encoded and stored in the corresponding position of the data warehouse table, thus ensuring the integrity and accuracy of the data. This process involves the precise management of timestamps and the accurate implementation of data encoding logic, enabling the data warehouse to effectively record and manage the historical status changes of target data over time, providing a reliable data basis for subsequent data analysis and decision support.

[0042] Step S104, if the sequence length corresponding to the start position is less than or equal to the current sequence length of the target data, convert the status change data into binary form and fill the conversion result into the range between the start position and the end position to obtain the updated data warehouse table.

[0043] Specifically, after determining the start position of the status change data in the data warehouse table, it is necessary to check whether the sequence length corresponding to this start position is less than or equal to the current sequence length of the target data. Here, the "sequence length" refers to the number of bits in the bit sequence used to store the status change information. If the sequence length of the start position meets the condition, the status change data will be converted into binary form. Binary form is a data form represented by 0s and 1s, which can store information in a very compact way. The conversion process usually involves mapping the enumerated values or identifiers of the status change data to a binary sequence. Once the conversion is completed, the obtained binary form will be filled into the corresponding area of the data warehouse table from the start position to the end position. This filling process ensures that the status change information is accurately recorded in the specified position of the data warehouse table, thus updating the content of the data warehouse table. In this way, the data warehouse can efficiently store and manage the historical status changes of target data over time, providing a solid foundation for subsequent data analysis and queries.

[0044] The present disclosure provides a data processing method, apparatus, device, and storage medium. By creating a data warehouse table containing target data with slowly changing states over time, the present disclosure achieves efficient data management. When a write request for state change data of the target data is received, the start and end positions of data writing are accurately determined based on the difference between the set time and the start time. Subsequently, the state change data is efficiently converted into binary form and filled into the corresponding positions. This binary representation is more compact than traditional storage methods, significantly reducing the requirement for data storage space. In addition, by directly locating and filling data, unnecessary data processing steps are avoided, thereby improving the speed and accuracy of data update. This method not only simplifies the data management process but also provides a more reliable basis for data analysis and decision support, ensuring the efficiency and accuracy of data processing.

[0045] In some alternative embodiments, determining the start position and end position of the state change data to be written based on the difference between the set time and the start time includes:

[0046] Obtaining the number of all state values of the target data to obtain the state quantity;

[0047] Performing a logarithm calculation on the state quantity to obtain the state bit number;

[0048] Calculating the difference between the set time and the start time to obtain a first difference;

[0049] Determining the start position and end position of the state change data to be written based on the state bit number and the first difference.

[0050] Specifically, obtaining the number of all state values of the target data means counting the total number of different states that the target data may have. For example, the inventory status of a commodity may have three states: "in stock", "sold", and "out of stock", so the state quantity is 3. The state quantity is the basis for determining the state bit number, and the state bit number is obtained by performing a logarithm calculation on the state quantity. Usually, the logarithm is taken to the base 2 and rounded up to ensure that all state values can be represented with the fewest number of bits, that is, satisfying the formula where n represents the state quantity and m represents the state bit number. For example, when the state quantity is 3, the logarithm calculation result is 2 (because 2 squared is greater than or equal to 3), which means that 2 bits in binary form are required to represent all states. Then, calculating the difference between the set time and the start time (the first difference) is to determine the specific position of the state change data in the bit sequence.

[0051] For the start time and the set time, the start time is the start time of the status change record, while the set time is the time point at which the status change needs to be written currently. The first difference represents the time interval elapsed from the start time to the set time. Based on the number of status bits and the first difference, the start position and the end position of the status change data in the bit sequence can be accurately calculated. The start position is the position where the status change data starts to be written, and the end position is the position where the writing ends. In this way, the status change data can be accurately encoded into the bit sequence, and then efficient storage and management of the data can be achieved, providing support for subsequent data analysis and query.

[0052] In this way, by obtaining the number of all status values of the target data and performing a logarithmic calculation to determine the number of status bits, and calculating the difference between the set time and the start time to determine the writing position of the status change data, the status change information can be efficiently encoded and stored with the smallest storage space. This method not only reduces the redundancy of data storage and improves the storage efficiency. In addition, since this solution can directly locate the status change data at a specific time point, the reading and query of the data can be made faster and more accurate. Moreover, this precise encoding and positioning mechanism also simplifies the data management process, reduces the complexity of data processing, and thus provides a more reliable and efficient data basis for data analysis and decision support.

[0053] In some alternative embodiments, determining the start position and the end position where the status change data is written according to the number of status bits and the first difference includes:

[0054] Calculating the product of the first difference and the number of status bits, and taking the sequence length corresponding to the product result as the start position;

[0055] Accumulating the sequence length corresponding to the start position and the number of status bits, and taking the sequence length corresponding to the accumulation result as the end position.

[0056] Specifically, calculating the product of the first difference and the number of status bits is to determine the specific start position of the status change data in the bit sequence. The first difference refers to the difference between the set time and the start time, which reflects the time interval elapsed from the start of the status change to the current set time. The number of status bits is obtained by performing a logarithmic calculation on the number of all possible status values of the target data, indicating the number of binary bits required for each status.

[0057] Multiply the first difference by the number of status bits. The resulting value represents the total number of status bits required from the start time to the set time. The sequence length corresponding to this total number is the starting position of the status change data. Next, to determine the ending position, add the sequence length corresponding to the starting position to the number of status bits. The result of the addition represents the space required starting from the starting position and adding one more status bit. The sequence length corresponding to this addition result is the ending position. In this way, it is possible to accurately locate the writing interval of the status change data in the bit sequence, and thus ensure the accurate storage of data and subsequent efficient query. This method effectively optimizes the data storage structure and improves the efficiency and accuracy of data processing.

[0058] In this way, the starting position is determined by calculating the product between the first difference and the number of status bits, and the ending position is obtained by adding the starting position to the number of status bits. This solution can accurately locate the storage interval of the status change data in the bit sequence. This method optimizes the data storage structure, making the writing and querying of data more efficient and accurate, reducing unnecessary data processing steps, and thus improving the speed of data processing and the overall performance of the system.

[0059] In some alternative embodiments, the method further includes:

[0060] If the sequence length corresponding to the starting position is greater than the current sequence length of the target data, then fill the interval between the end position of the target data and the starting position with a preset value.

[0061] Specifically, the current sequence length of the target data refers to the number of bits actually occupied by the target data in the bit sequence. When the sequence length corresponding to the starting position is greater than the current sequence length of the target data, it means that there is an idle area in the bit sequence, and this part of the area needs to be filled to maintain the integrity and consistency of the data. The preset value is usually a default binary value, such as 0, used to fill these idle positions. The specific implementation process of this solution includes: detecting the bit difference between the starting position and the end position of the target data, and then filling this difference area with the preset value. This filling operation ensures the continuity and integrity of the bit sequence, avoids gaps between data, and thus improves the efficiency of data storage and the accuracy of query. In this way, it is possible to effectively manage the data layout in the bit sequence and ensure the compact storage and fast access of data.

[0062] In this way, by filling the idle area between the end position and the start position of the target data with a preset value, the continuity and integrity of the bit sequence can be ensured, avoiding gaps between data. The filling operation of this solution makes the bit sequence more compact, reduces waste of storage space, and effectively improves the efficiency of data storage. At the same time, this method also simplifies the data management process. Since the filled bit sequence is easier to perform subsequent data processing and query operations, the overall performance of the system and the accuracy of data processing are improved.

[0063] To facilitate the overall understanding of the data writing operation of the data warehouse table of the present disclosure, the following is an example. Refer to Figure 2 , Figure 2 which is a schematic flowchart of writing data into the data warehouse table in an embodiment of the present disclosure. Specifically, first, obtain the data warehouse table and the start time when the target data status changes. Among them, the data warehouse table contains target data (also called col_s) whose status changes slowly over time. Assume that the status values corresponding to col_s include enum1, enum2... enumn-1, and obtain data A = [enum0, enum1, enum2... enumn-1], where enum0 is a special status value used as the default status value when the change is discontinuous. Take the logarithm of n to the base 2 to get where n represents the number of statuses and m represents the number of status bits.

[0064] The flowchart includes the following steps:

[0065] Step S201, obtain the length (l) of col_s, that is, first need to know the current bit sequence length of col_s.

[0066] Step S202, calculate the time difference △T, and determine the start position s_i and the end position e_i: Calculate the time difference (△t) between the specified time (t) and the start time (s_t). This time difference is used to determine the position of the status change in the bit sequence. Then, according to the calculated time difference (△t), determine the start position (s_i) and the end position (e_i) of the status change data in the bit sequence. Among them, the start position is obtained by multiplying the time difference by the number of bits (m) required for each status, and the end position is the start position plus the number of bits (m).

[0067] Step S203, determine whether the calculated start position (s_i) is greater than the length (l) of the current bit sequence.

[0068] Step S204, if so, it indicates that the bit sequence needs to be padded to ensure there is enough space to store the new state change data. If the starting position (s_i) is greater than the length (l) of the current bit sequence, fill the positions from l to s_i of col_s with 0 bits, that is, fill 0 bits at the end of the bit sequence until the length reaches s_i.

[0069] Step S205, if the starting position (s_i) is not greater than the length (l) of the current bit sequence, obtain the index (x) of enumx in the array A, that is, obtain the index (x) of the state value corresponding to the specified time (t) in the enumeration array A.

[0070] Step S206, encode the index x into an m-bit binary fragment. This binary fragment represents the new state value.

[0071] Step S207, insert the encoded binary fragment into the positions from s_i to e_i of the bit sequence.

[0072] Through this process, the new state change is recorded in col_s. This method can efficiently update the state information in the slowly changing dimension table while maintaining compact data storage and fast query capabilities. This method is particularly suitable for processing data that changes over time but with a low frequency, such as user status, product classification, etc.

[0073] See Figure 3 , Figure 3 FIG. is a schematic diagram of inserting an encoded fragment into a bit sequence in an embodiment of the present disclosure. This flowchart shows the working principle of converting a decimal value into a binary code and inserting it into an existing bit sequence, which is consistent with the process of encoding and updating the slowly changing dimension table in the technical solution of this application. Assume there is an existing bit sequence, which is used to store the historical state of slowly changing dimension data. According to the requirements of the business logic or data model, determine the starting position (s_i) and ending position (e_i) where the new state value needs to be inserted into the bit sequence. At this time, convert the state value to be inserted (in this example, the decimal number 5) into binary format. In this example, the binary representation of 5 is 0101. According to the determined number of bits (m), calculate the binary fragment to be inserted. In this example, if each state value is represented by 3-bit binary numbers, then the binary code of 5 is 0101. Insert the calculated binary fragment (0101) into the specified position of the bit sequence (from s_i to e_i). This may involve moving some of the existing data in the bit sequence to make room, or if the length of the bit sequence is insufficient, it may be necessary to first expand the bit sequence. The bit sequence is updated to contain the new state value. This updated bit sequence can now be used for subsequent data query and analysis.

[0074] In the technical solution of this application, this process is automated and implemented through user-defined functions (UDFs). The UDF can encapsulate this conversion and insertion logic, enabling the efficient update of the status information in the data warehouse table when dealing with slowly changing dimension data. This method reduces data redundancy, optimizes the use of storage space, and improves the efficiency of data querying.

[0075] In some alternative embodiments, the method further includes:

[0076] In response to a query request for status data corresponding to a target time for a target data query, determining a starting position and an ending position corresponding to the status data of the target time according to the difference between the target time and the starting time;

[0077] If the sequence length of the starting position is greater than or equal to a preset value, or the sequence length of the ending position is less than or equal to the current sequence length of the target data, obtaining a bit sequence segment between the starting position and the ending position, and converting the bit sequence segment into a decimal form to obtain the status data corresponding to the target data.

[0078] Specifically, when a query request is received with the aim of retrieving the status data of a target data at a specific target time, the time difference between the target time and the starting time of the data record is first calculated. This difference is used to determine the specific positions of the status data of the target time in the storage structure, namely the starting position and the ending position. The starting position refers to the position in the bit sequence where the status data starts to be retrieved, while the ending position is the position where the retrieval ends. The preset value is a predefined threshold used to ensure the effectiveness of the query operation, and it may be related to the storage density of the data or the granularity of the query. The current sequence length of the target data refers to the actual number of bits occupied by the target data in the bit sequence.

[0079] If the sequence length at the starting position is greater than or equal to the preset value, or the sequence length at the ending position is less than or equal to the current sequence length of the target data, this indicates that there is sufficient data within the specified time range for an effective query. In this case, the segment between the starting position and the ending position will be extracted from the bit sequence. The bit sequence segment is a piece of binary-encoded data that contains the status information of the target data at the target time. Subsequently, this bit sequence segment is converted into decimal form, and this conversion process involves decoding the binary number into a decimal number to obtain the specific status data corresponding to the target data at the target time. Since this method can directly locate the relevant data segment, it improves the efficiency of data retrieval. And by obtaining the status information based on precise time points, it also improves the accuracy of the query. In addition, by using binary sequences to store and retrieve data, this solution also optimizes the use of storage space, enabling the data warehouse to more effectively manage a large amount of slowly changing dimensional data. This method is particularly suitable for processing data that changes over time but with a low frequency, such as user status, product classification, etc., providing a reliable data basis for data analysis and decision support.

[0080] In this way, by precisely calculating the difference between the target time and the starting time to determine the starting and ending positions of the status data in the bit sequence, and extracting and converting the bit sequence segment between these positions into decimal form when specific conditions are met. This solution can efficiently and accurately retrieve the status information of the target data at a specific time point. This method not only improves the efficiency and accuracy of the query, reduces unnecessary data processing, but also optimizes the use of storage space because it only stores the necessary status change data. In addition, by directly extracting information from the compact bit sequence, this solution also speeds up the data retrieval speed, providing a more reliable basis for data analysis and decision support, thus significantly enhancing the overall performance and reliability of data processing.

[0081] In some alternative embodiments, the method further includes:

[0082] If the sequence length at the starting position is less than the preset value, or the sequence length at the ending position is greater than the current sequence length of the target data, then a null value is returned.

[0083] Specifically, when processing a query request, first determine the starting position and ending position corresponding to the status change data of the query based on the set time and the starting time. The starting position and the ending position respectively represent the positions where the status change data starts and ends in the bit sequence. The preset value is a set threshold used to determine whether the sequence length at the starting position is large enough to contain valid status change information. The current sequence length of the target data refers to the number of bits actually occupied by the target data in the bit sequence. If the sequence length at the starting position is less than the preset value, it indicates that there is not enough data before this position to represent a complete status change; or if the sequence length at the ending position is greater than the current sequence length of the target data, it means that the query range exceeds the actual storage range of the target data. In this case, a null value will be returned, indicating that valid status change data cannot be obtained. This processing mechanism ensures the accuracy and reliability of the query results, avoids the return of incorrect or incomplete data, and thus improves the quality of data query and the stability of the system.

[0084] In this way, by returning a null value when the sequence length at the starting position is less than the preset value or the sequence length at the ending position is greater than the current sequence length of the target data, this solution can effectively avoid returning incomplete or invalid status change data, thereby ensuring the accuracy and reliability of the query results. This method improves the quality of data query, prevents the impact of incorrect data on analysis and decision-making, and enhances the stability of the system and the accuracy of data processing.

[0085] In some alternative embodiments, determining the starting position and ending position corresponding to the status data of the target time according to the difference between the target time and the starting time includes:

[0086] Obtain the number of all status values of the target data to get the status quantity;

[0087] Perform a logarithmic calculation on the status quantity to obtain the status bits;

[0088] Calculate the difference between the target time and the starting time to get the second difference;

[0089] Calculate the product of the second difference and the status bits, and use the sequence length corresponding to the product result as the starting position;

[0090] Accumulate the sequence length corresponding to the starting position and the status bits, and use the sequence length corresponding to the accumulation result as the ending position.

[0091] Specifically, first, it is necessary to obtain the number of all possible state values of the target data, which is called the "number of states". The number of states refers to the total number of all different states that the target data may change in the business logic. For example, the inventory status of a product may include "in stock", "sold", and "out of stock", etc. Next, a logarithmic calculation is performed on the number of states to determine the minimum number of bits required to represent these states, and this number of bits is called the "number of state bits". The calculation of the number of state bits is usually based on the number of bits of the smallest binary number that can cover all state values. For example, if there are 3 states, at least 2 bits of binary numbers are required (because 2 to the power of 2 is equal to 4, which is sufficient to represent 3 different states).

[0092] Then, calculate the difference between the target time and the start time, and this difference is called the "second difference". The second difference reflects the number of time units elapsed from the start of the record to the target time, and is used to determine the relative position of the state change data in the bit sequence. Next, calculate the product of the second difference and the number of state bits, and the sequence length corresponding to this product result is used as the "starting position" in the bit sequence. The starting position is the position where the state change data starts to be queried in the bit sequence.

[0093] Finally, add the sequence length corresponding to the starting position and the number of state bits, and the sequence length corresponding to the resulting sum is used as the "ending position" in the bit sequence. The ending position is the position where the state change data ends to be queried in the bit sequence. In this way, it is possible to accurately locate the state change data in the bit sequence, and then efficient data storage and query can be achieved. This method not only optimizes the data storage structure and reduces the demand for storage space, but also improves the efficiency and accuracy of data processing.

[0094] In this way, by obtaining the number of all state values of the target data and calculating the number of state bits, and then combining the difference between the target time and the start time (the second difference) to determine the starting and ending positions of the state change data, this solution realizes the accurate positioning and efficient management of slowly changing dimension data. This method not only optimizes the data storage structure and reduces the demand for storage space, but also improves the efficiency of data update and query by precisely controlling the writing and reading positions of the data. In addition, it simplifies the data management process because the data state at a specific time point can be directly located, thus providing more accurate and reliable information for data analysis and decision support.

[0095] In some alternative embodiments, the method further includes:

[0096] After obtaining the state change data of the target data at the set time, merge the state change data of the previous time unit with the state change data of the set time to obtain the latest partition data;

[0097] Delete the status data of the previous time unit and save the latest partition data.

[0098] Specifically, after obtaining the status change data of the set time target data, merge the status change data of the previous time unit with the status change data of the set time. Here, the "previous time unit" usually refers to the nearest time point before the set time, such as the previous day or the previous hour. The merging process involves integrating the status change information of the two time points to form a complete data set, which is called the latest partition data. The latest partition data contains all the status change information from the start time to the set time, ensuring the continuity and integrity of the data. After the merging is completed, the status data of the previous time unit will be deleted to save storage space and optimize data management. The deletion operation is usually carried out after confirming that the latest partition data has been correctly generated and saved to avoid data loss. In this way, the data in the data warehouse can be effectively managed and updated, ensuring the timeliness and accuracy of the data, and providing a solid foundation for subsequent data analysis and decision support.

[0099] In this way, by merging the status change data of the previous time unit with the status change data of the set time, obtaining the latest partition data, and deleting the status data of the previous time unit, this solution can effectively update the information in the data warehouse, ensuring the timeliness and accuracy of the data. This method reduces the occupancy of storage space because only the latest data partition is retained, thereby improving the efficiency of data management. At the same time, it simplifies the data query process because only the latest partition data needs to be accessed during query, thereby accelerating the query speed and enhancing the overall performance of the system.

[0100] To facilitate the overall understanding of the data query operation of the data warehouse table of the present disclosure, an example is given as follows. See Figure 4 , Figure 4 is a schematic flowchart of the data warehouse table data query in the embodiment of the present disclosure. This flowchart describes the process for decoding slowly changing dimension data in the technical solution of the present application. The flowchart includes the following steps:

[0101] Step S401, obtain the length (l) of col_s: First, obtain the current length of col_s, which is to determine the total size of the bit sequence for subsequent operations.

[0102] Step S402, calculate the time difference (△t), and determine the starting position s_i and the ending position e_i: Calculate the time difference (△t) between the target time (t) and the starting time (s_t), and determine the starting position (s_i) and the ending position (e_i) of the status change data in the bit sequence according to this time difference.

[0103] Step S403, determine whether the calculated starting position (s_i) is less than 0, or whether the ending position (e_i) is greater than or equal to the length (l) of the bit sequence.

[0104] Step S404, if either condition is satisfied, it means that the specified time is not within the recording range of the data, so return a null value.

[0105] Step S405, if the condition is not satisfied, obtain the bit sequence segment from s_i to e_i from col_s.

[0106] Step S406, convert the extracted binary segment into a decimal number x, and this decimal number represents the index of the state value at the specified time point.

[0107] Step S407, use the converted decimal number (x) as an index to obtain the corresponding state value (enum_x) from the state enumeration array A, and this state value is the dimensional state at the specified time point.

[0108] Through this process, the state value at any time point can be efficiently queried from the slowly changing dimensional data. This method not only reduces the storage requirements but also improves the query efficiency because it directly locates the relevant segment in the bit sequence without having to scan the entire data set. In addition, by using bit sequences and binary encoding, the storage of data becomes more compact, thus optimizing the use of storage space.

[0109] See Figure 5 , Figure 5 is a schematic diagram of obtaining a decoded segment from a bit sequence in an embodiment of the present disclosure. This flowchart shows the process of extracting a specific segment from a bit sequence and converting it into a decimal value, which is typically used to decode the state value of slowly changing dimensional data. For example, determine the starting position (s_i) and ending position (e_i) of the segment to be extracted from the bit sequence. These positions are calculated based on the time difference and reflect the storage position of the state change in the bit sequence. Extract the segment from s_i to e_i from the bit sequence. In this example, the extracted segment is "0100". Convert the extracted binary segment "0100" into a decimal value. In this example, "0100" is converted into a decimal value of 4. Use the converted decimal value as an index to obtain the corresponding state value from the predefined state enumeration array A. In this example, if the state value corresponding to index 4 in array A is "out of stock", the decoding process will return "out of stock" as the result.

[0110] This process is the core part of decoding UDF (User-Defined Function) in the technical solution of this application. It can quickly and accurately retrieve the dimension status at any time point from the compactly stored bit sequence. This method not only improves the efficiency of data storage but also speeds up the query speed because it directly locates the relevant fragments in the bit sequence without scanning the entire data set. In addition, by using bit sequences and binary encoding, the storage of data becomes more compact, thus optimizing the use of storage space.

[0111] The following describes the device embodiments of this application, which can be used to execute the data processing method in the above embodiments of this application. For the details not disclosed in the device embodiments of this application, please refer to the embodiments of the above data processing method of this application.

[0112] The present disclosure also provides a data processing device 600, as Figure 6 shown, including:

[0113] A first acquisition module 601, configured to acquire a data warehouse table, where the data warehouse table contains target data with slow state changes over time;

[0114] A second acquisition module 602, configured to acquire the starting time when the state of the target data changes;

[0115] A response module 603, configured to, in response to a request for writing state change data for the target data, determine the starting position and the ending position for writing the state change data according to the difference between the set time and the starting time;

[0116] An update module 604, configured to, if the sequence length corresponding to the starting position is less than or equal to the current sequence length of the target data, convert the state change data into a binary form and fill the converted result into the range between the starting position and the ending position to obtain an updated data warehouse table.

[0117] In some alternative embodiments, the response module 603 determines the starting position and the ending position for writing the state change data according to the difference between the set time and the starting time, including;

[0118] Acquire the number of all state values of the target data to obtain the state quantity;

[0119] Perform a logarithmic calculation on the state quantity to obtain the number of state bits;

[0120] Calculate the difference between the set time and the starting time to obtain a first difference;

[0121] Determine the starting position and the ending position for writing the state change data according to the number of state bits and the first difference.

[0122] In some alternative embodiments, the response module 603 determines the starting position and the ending position for writing the status change data according to the number of status bits and the first difference, including:

[0123] Calculate the product of the first difference and the number of status bits, and use the sequence length corresponding to the product result as the starting position;

[0124] Accumulate the sequence length corresponding to the starting position and the number of status bits, and use the sequence length corresponding to the accumulation result as the ending position.

[0125] In some alternative embodiments, the update module 604 is further configured to, if the sequence length corresponding to the starting position is greater than the current sequence length of the target data, fill the range between the ending position and the starting position of the target data with a preset value.

[0126] In some alternative embodiments, the apparatus further includes a query module, configured to, in response to a query request for querying the status data corresponding to the target time for the target data, determine the starting position and the ending position corresponding to the status data of the target time according to the difference between the target time and the starting time;

[0127] If the sequence length of the starting position is greater than or equal to a preset value, or the sequence length of the ending position is less than or equal to the current sequence length of the target data, obtain the bit sequence segment between the starting position and the ending position, and convert the bit sequence segment into a decimal form to obtain the status data corresponding to the target data.

[0128] In some alternative embodiments, the query module determines the starting position and the ending position corresponding to the status data of the target time according to the difference between the target time and the starting time, including:

[0129] Obtain the number of all status values of the target data to obtain the number of statuses;

[0130] Perform a logarithmic calculation on the number of statuses to obtain the number of status bits;

[0131] Calculate the difference between the target time and the starting time to obtain a second difference;

[0132] Calculate the product of the second difference and the number of status bits, and use the sequence length corresponding to the product result as the starting position;

[0133] Accumulate the sequence length corresponding to the starting position and the number of status bits, and use the sequence length corresponding to the accumulation result as the ending position.

[0134] In some alternative embodiments, the apparatus further includes a merging module, configured to merge the state change data of the previous time unit with the state change data of the set time after obtaining the state change data of the set time target data, so as to obtain the latest partition data;

[0135] Delete the state data of the previous time unit and save the latest partition data.

[0136] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0137] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0138] Figure 7 FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0139] As Figure 7 shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0140] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0141] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the small program distribution described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the data processing method in any other suitable manner (e.g., by means of firmware).

[0142] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0145] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0146] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0147] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0148] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0149] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A data processing method, the method comprising: Obtaining a data warehouse table, wherein the data warehouse table contains target data whose state changes slowly over time; Obtaining the starting time when the target data state changes; In response to a state change data write request for the target data, determining a start point position and an end point position of the state change data write according to a difference between a set time and the start time; If the sequence length corresponding to the starting position is less than or equal to the current sequence length of the target data, the state change data is converted into binary form, and the conversion result is filled into the interval between the starting position and the end position to obtain an updated data warehouse table.

2. The method according to claim 1, wherein: The step of determining the starting point and the ending point of writing the state change data according to the difference between the set time and the start time includes: Obtain the number of all state values ​​of the target data to obtain the state quantity; Performing logarithmic calculation on the number of states to obtain the number of state bits; Calculating the difference between the set time and the start time to obtain a first difference; The starting position and the ending position of writing the state change data are determined according to the number of state bits and the first difference.

3. The method according to claim 2, wherein: The step of determining the starting position and the ending position of writing the state change data according to the state bit number and the first difference value includes: Calculate the product of the first difference and the number of state bits, and use the sequence length corresponding to the product result as the starting point; The sequence length corresponding to the starting position and the number of state bits are accumulated, and the sequence length corresponding to the accumulation result is used as the end position.

4. The method according to claim 1, wherein: The method further comprises: If the sequence length corresponding to the starting position is greater than the current sequence length of the target data, the interval between the end position of the target data and the starting position is filled with a preset value.

5. The method according to any one of claims 1 to 4, wherein: The method further comprises: In response to a query request for querying the status data corresponding to the target time for the target data, determining the starting position and the ending position corresponding to the status data of the target time according to the difference between the target time and the start time; If the sequence length of the starting position is greater than or equal to a preset value, or the sequence length of the end position is less than or equal to the current sequence length of the target data, then the bit sequence fragment between the starting position and the end position is obtained, and the bit sequence fragment is converted into decimal form to obtain the status data corresponding to the target data.

6. The method according to claim 5, wherein: The step of determining the starting point position and the ending point position corresponding to the state data of the target time according to the difference between the target time and the start time includes: Obtain the number of all state values ​​of the target data to obtain the state quantity; Performing logarithmic calculation on the number of states to obtain the number of state bits; Calculating the difference between the target time and the start time to obtain a second difference; Calculate the product of the second difference and the number of state bits, and use the sequence length corresponding to the product result as the starting point; The sequence length corresponding to the starting position and the number of state bits are accumulated, and the sequence length corresponding to the accumulation result is used as the end position.

7. The method according to any one of claims 1 to 4, wherein: The method further comprises: After obtaining the state change data of the target data at the set time, merging the state change data of the previous time unit with the state change data of the set time to obtain the latest partition data; The state data of the previous time unit is deleted, and the latest partition data is saved.

8. A data processing device, comprising: A first acquisition module is used to acquire a data warehouse table, wherein the data warehouse table contains target data whose state changes slowly over time; A second acquisition module is used to acquire the start time when the target data state changes; A response module, configured to respond to a state change data write request for the target data, and determine a start point position and an end point position of the state change data write according to a difference between a set time and the start time; An updating module is used to convert the state change data into binary form if the sequence length corresponding to the starting position is less than or equal to the current sequence length of the target data, and fill the conversion result into the interval range between the starting position and the end position to obtain an updated data warehouse table.

9. The device according to claim 8, wherein: The response module determines the starting position and the ending position of the state change data writing according to the difference between the set time and the start time, including: Obtain the number of all state values ​​of the target data to obtain the state quantity; Performing logarithmic calculation on the number of states to obtain the number of state bits; Calculating the difference between the set time and the start time to obtain a first difference; The starting position and the ending position of writing the state change data are determined according to the number of state bits and the first difference.

10. The device according to claim 9, wherein: The response module determines the starting position and the ending position of the state change data writing according to the state bit number and the first difference value, including: Calculate the product of the first difference and the number of state bits, and use the sequence length corresponding to the product result as the starting point; The sequence length corresponding to the starting position and the number of state bits are accumulated, and the sequence length corresponding to the accumulation result is used as the end position.

11. The device according to claim 8, wherein: The updating module is further configured to fill the interval between the end position of the target data and the starting position with a preset value if the sequence length corresponding to the starting position is greater than the current sequence length of the target data.

12. The device according to any one of claims 8 to 11, wherein: The device further comprises a query module, which is used to respond to a query request for querying the state data corresponding to the target time for the target data, and determine the starting position and the ending position corresponding to the state data of the target time according to the difference between the target time and the start time; If the sequence length of the starting position is greater than or equal to a preset value, or the sequence length of the end position is less than or equal to the current sequence length of the target data, then the bit sequence fragment between the starting position and the end position is obtained, and the bit sequence fragment is converted into decimal form to obtain the status data corresponding to the target data.

13. The device according to claim 12, wherein: The query module determines the starting position and the ending position corresponding to the status data of the target time according to the difference between the target time and the start time, including: Obtain the number of all state values ​​of the target data to obtain the state quantity; Performing logarithmic calculation on the number of states to obtain the number of state bits; Calculating the difference between the target time and the start time to obtain a second difference; Calculate the product of the second difference and the number of state bits, and use the sequence length corresponding to the product result as the starting point; The sequence length corresponding to the starting position and the number of state bits are accumulated, and the sequence length corresponding to the accumulation result is used as the end position.

14. The device according to any one of claims 8 to 11, wherein: The device further comprises a merging module, which is used to merge the state change data of the previous time unit with the state change data of the set time after obtaining the state change data of the target data at the set time, so as to obtain the latest partition data; The state data of the previous time unit is deleted, and the latest partition data is saved.

15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.