Data table generation method and device
By statistically and analyzing the frequency of field occurrence of business data tables in the data warehouse, and using word frequency-inverse document frequency technology to determine the target field as the dimension of the data model, the problems of field redundancy and resource waste in the data model design are solved, and more efficient and flexible data table generation is achieved.
Patent Information
- Application Number
- CN202210293153.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-03-23
AI Technical Summary
In the data warehouse, during the design process of real-time DWD data model, it is easy to cause redundancy in model fields when confirming dimensions, resulting in waste of computing resources, and to reduce the redundancy cost and is not easy to be comprehensive by collecting customer needs or business experience, which increases the risk of model optimization.
By determining the row record information of the business data table and the target data table associated with the business topic based on the business topic, the frequency of occurrence of each field in each business data table is counted, and the target field with the statistical value meets the preset conditions is obtained using the word frequency-inverse document frequency technology, and it is used as the dimension of the target data table to generate each dimension table, select dimension attributes, determine attribute values, and generate target data tables.
It effectively reduces model field redundancy, saves computing resources, reduces the cost of collecting customer needs, improves model coverage and flexibility, and reduces the risk of model optimization.
Smart Images

Figure CN114691682B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, specifically to the field of data processing technology, and more particularly to a method and device for generating a data table. Background Art
[0002] In the data warehouse, the detail data layer DWD serves as an isolation layer between the business layer and the data warehouse, and is mainly used to clean and normalize data for the operational data storage system ODS. The current design process of the real-time DWD data model includes: (1) selecting the business subject; (2) confirming the granularity, which refers to the degree of business detail expressed by a record in the fact table; (3) confirming the dimension: selecting the dimensional information that can clearly describe the environment in which the business process is located; (4) confirming the facts: determining the various indicators to be measured by the data model. However, when confirming the dimension process, the usual practice is to cover all the source data fields, which can easily cause model field redundancy and waste of computing resources. In order to reduce model redundancy, it is necessary to collect a large amount of customer needs, or to extract high-frequency fields of source data as model fields based on past business experience. However, this method has a high cost for collecting needs and is not easy to collect comprehensively, which increases the risk of later model optimization. Summary of the invention
[0003] The embodiments of the present disclosure provide a method and device for generating a data table.
[0004] In a first aspect, an embodiment of the present disclosure provides a data table generation method, comprising: determining at least one business data table associated with the business subject and row record information of a target data table according to the business subject; performing statistics on the occurrence frequency of each field in each business data table to obtain at least one target field whose statistical value meets a preset condition, and using each target field as each dimension of the target data table to generate each dimension table; selecting dimension attributes in each dimension table; determining the attribute value of each dimension based on the row record information of the target data table and the attributes of each dimension, and generating a target data table that at least includes the attribute value of each dimension.
[0005] In some embodiments, the frequency of occurrence of each field in each business data table is counted to obtain at least one target field whose statistical value meets preset conditions, including: using the word frequency-inverse document frequency technology to count the frequency of occurrence of each field in the script library corresponding to each business data table to obtain at least one target field whose statistical value meets preset conditions, the statistical value represents the product of the word frequency of the field and the inverse document frequency of the corresponding field.
[0006] In some embodiments, the preset condition is that the statistical value is greater than a threshold value and / or the statistical value is located before a preset sequence number after all statistical values are sorted; the word frequency-inverse document frequency technology is used to count the occurrence frequency of each field in the script library corresponding to each business data table, and at least one target field whose statistical value meets the preset condition is obtained, including: according to each preset condition in multiple preset conditions, each field in the script library is divided to obtain each type of field corresponding to each preset condition; the word frequency-inverse document frequency technology is used to count the occurrence frequency of each type of field to obtain at least one sub-target field of each type of field whose statistical value meets the corresponding preset condition; the sub-target fields of each type of field are merged to obtain each target field.
[0007] In some embodiments, before using the word frequency-inverse document frequency technology to count the occurrence frequency of each field in each business data table in the script library to obtain at least one target field whose statistical value meets preset conditions, it also includes: marking each field in the script library with stop words so that the fields marked as stop words are ignored in the subsequent statistics of each field in the script library.
[0008] In some embodiments, each preset condition is set based on each type of field information; each preset condition corresponds to a type of field information.
[0009] In some embodiments, before selecting the dimension attributes in each dimension table, it also includes: determining the associated fields related to each dimension according to the dimensions of the target data table; using the associated fields as the dimensions of the target data table, and merging them with the existing dimensions of the target data table to generate the dimensions of the final target data table.
[0010] In some embodiments, determining the row record information of the target data table includes: analyzing the business data table, and selecting the information with the finest detail representing the business subject as the row record information of the target data table.
[0011] In some embodiments, after determining the row record information of the target data table, the method further includes: using the row record information of the target data table as the primary key of the target data table.
[0012] In some embodiments, the method further includes: displaying the target data table.
[0013] In a second aspect, an embodiment of the present disclosure provides a data table generating device, comprising: a first determining unit, configured to determine at least one business data table associated with a business subject and row record information of a target data table according to the business subject; a statistical unit, configured to count the occurrence frequency of each field in each business data table, obtain at least one target field whose statistical value satisfies a preset condition, and use each target field as a dimension of the target data table to generate each dimension table; a selecting unit, configured to select dimension attributes in each dimension table; a first generating unit, configured to determine the attribute value of each dimension based on the row record information of the target data table and the attributes of each dimension, and generate a target data table including at least the attribute value of each dimension.
[0014] In some embodiments, the statistical unit is further configured to use the word frequency-inverse document frequency technology to count the occurrence frequency of each field in each business data table in the script library corresponding to each business data table, and obtain at least one target field whose statistical value meets the preset conditions, and the statistical value represents the product of the word frequency of the field and the inverse document frequency of the corresponding field.
[0015] In some embodiments, the apparatus further comprises: a marking unit configured to mark each field in the script library with stop words, so that each field marked as a stop word is ignored in subsequent statistics of each field in the script library.
[0016] In some embodiments, the preset conditions in the statistical unit are that the statistical value is greater than a threshold and / or the statistical value is located before a preset sequence number after all statistical values are sorted; the statistical unit includes: a field division module, configured to divide the fields in the script library according to each preset condition in a plurality of preset conditions, and obtain each type of field corresponding to each preset condition; a frequency statistics module, configured to use the word frequency-inverse document frequency technology to count the occurrence frequency of each type of field, and obtain at least one sub-target field of each type of field whose statistical value meets the corresponding preset condition; a merging module, configured to merge the sub-target fields of each type of field to obtain each target field.
[0017] In some embodiments, each preset condition in the statistical unit is set based on each type of field information; each preset condition in the statistical unit corresponds to a type of field information.
[0018] In some embodiments, the device also includes: a second determination unit, configured to determine the associated fields related to each dimension based on the dimensions of the target data table; a second generation unit, configured to use the associated fields as the dimensions of the target data table, and merge them with the existing dimensions of the target data table to generate the dimensions of the final target data table.
[0019] In some embodiments, the first determining unit is further configured to analyze the business data table and select information representing the business subject with the finest level of detail as row record information of the target data table.
[0020] In some embodiments, the apparatus further includes: a setting unit configured to use the row record information of the target data table as the primary key of the target data table.
[0021] In some embodiments, the device further includes: a display unit configured to display the target data table.
[0022] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0023] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0024] The data table generation method and device provided by the embodiments of the present disclosure, by determining at least one business data table associated with the business subject and the row record information of the target data table according to the business subject, counting the occurrence frequency of each field in each business data table, obtaining at least one target field whose statistical value meets the preset condition, and using each target field as each dimension of the target data table, generating each dimension table, selecting the dimension attribute in each dimension table, determining the attribute value of each dimension based on the row record information of the target data table and the attributes of each dimension, and generating a target data table including at least the attribute value of each dimension, thereby realizing a more comprehensive and effective data table generation method and device.
[0025] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0027] Figure 1 is an exemplary system architecture diagram in which some embodiments of the present disclosure may be applied;
[0028] Figure 2 is a flow chart of an embodiment of a data table generating method according to the present disclosure;
[0029] Figure 3 is a schematic diagram of an application scenario of the data table generation method according to the present disclosure;
[0030] Figure 4 is a flow chart of another embodiment of a method for generating a data table according to the present disclosure;
[0031] Figure 5 is a structural schematic diagram of an embodiment of a data table generating device according to the present disclosure;
[0032] Figure 6 It is a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION
[0033] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0034] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0035] Figure 1 An exemplary system architecture 100 is shown to which the data table generating method or data table generating apparatus according to the embodiments of the present disclosure can be applied.
[0036] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0037] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to send transmission requests or receive transmission data, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as shopping applications, pickup applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0038] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens that support information browsing, including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, desktop computers, etc. Terminal devices 101, 102, 103 can interact with servers through network 104 to obtain information, etc. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules for providing distributed services, for example, or it can be implemented as a single software or software module. No specific limitation is made here.
[0039] The server 105 may be an application server that provides various services, such as an application server that supports transmission requests of the terminal devices 101, 102, and 103. The application server may perform judgment and other processing on the received transmission request and other data, and feed back the processing result (such as the target data table) to the terminal device.
[0040] It should be noted that the server 105 can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0041] It should be noted that the operations performed by the server 105 may also be performed by other electronic devices.
[0042] It should be noted that the data table generating method provided in the embodiments of the present disclosure is generally executed by the server 105 , and the corresponding data table generating device is generally disposed in the server 105 .
[0043] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.
[0044] Continue to refer Figure 2 , shows a process 200 of an embodiment of a data table generation method according to the present disclosure. The data table generation method comprises the following steps:
[0045] Step 201: Determine row record information of at least one business data table and a target data table associated with the business subject according to the business subject.
[0046] In this embodiment, when the execution subject (for example Figure 1When the server (shown in FIG. 1 ) receives a data table generation request or runs to perform a data table generation operation, it can first select a business theme, that is, the specific content to be reflected in the target data table, and then use the data source of the operational data storage (ODS) layer to conduct data research according to the selected business theme, so as to determine at least one business data table associated with the business theme and the row record information of the target data table associated with the business theme, that is, determine the granularity of the table. Among them, the business theme can be at least one of the three business events, the state of the business event, and the business process including multiple business events, such as small and medium-sized piece waybills, business processes for successful transactions, transaction order processes, etc. The number of business data tables is related to the business theme. In general, it is necessary to determine all business data tables associated with the business theme to make the generation of the target data table more comprehensive and effective. For example, if the business theme is small and medium-sized piece waybills, the business data tables of the waybill system determined include the waybill master table and multiple waybill extension tables. Here, whether all business data tables or part of the business data tables associated with the business theme are determined according to the business theme is not specifically limited.
[0047] Granularity is the smallest unit of the number of data rows. It represents the degree of business detail expressed by a record in a table. It is used to determine the level of detail of the business represented by a row in a table (such as a fact table) and determines the scalability of the data table dimension. Before determining the dimensions and facts of a table, the granularity of the table must be determined first, and each dimension and fact of the table must be consistent with the defined granularity. Granularity is usually expressed through business descriptions. For example, the row record information of the target data table can be a waybill number, order number, etc. By clarifying the granularity of the table in advance, it can be ensured that there is no confusion about the meaning of the rows in the fact table, and that all facts are recorded at the same level of detail.
[0048] It should be noted that there cannot be multiple facts of different granularities in the same fact table. The granularity of all facts in the fact table needs to be consistent with the granularity declared in the table.
[0049] Step 202, count the occurrence frequency of each field in each business data table to obtain at least one target field whose statistical value meets the preset condition, and use each target field as each dimension of the target data table to generate each dimension table.
[0050] In this embodiment, the above-mentioned execution subject can count the frequency of occurrence of each field in each business data table determined in step 201, obtain the statistical value of each field, determine whether each statistical value meets the preset condition, and use the field whose statistical value meets the preset condition as the target field to obtain at least one target field, and use each target field as each dimension of the target data table to generate each dimension table. Among them, the target data table is used to reflect the business theme. The preset condition is pre-set based on selecting dimensional information that can clearly describe the environment in which the business process is located. Here, the same field in the target field is used as only one dimension of the target data table, avoiding the redundancy of the target data table dimension caused by field redundancy.
[0051] As an example, when the business subject is small and medium-sized parcel waybills, the target fields obtained after statistics may include waybill number, waybill status, waybill type, waybill identification, warehouse, distribution center, delivery method, address, sorting station, waybill weight, waybill volume, etc. The identified high-frequency fields are integrated into the model as dimensional information.
[0052] Step 203: Select dimension attributes in each dimension table.
[0053] In this embodiment, the above-mentioned execution entity can select the names of the members of each dimension from the corresponding dimension tables as dimension attributes according to the dimensions of the target data table determined in step 202 and the preset data selection principles. Among them, the dimension attributes may include some or all attributes related to the business subject, and there is a corresponding relationship between the attributes of each dimension and the row record information of the target data table. The attributes of each dimension can be used as the columns of the target data table. For the fact table, the attributes of each dimension are equivalent to the facts of the fact table. As an example, when the business subject is the order business process, the selected attributes of each dimension include product ID, product price, and purchase quantity.
[0054] Step 204 : Based on the row record information of the target data table and the attributes of each dimension, determine the attribute value of each dimension, and generate a target data table including at least the attribute value of each dimension.
[0055] In this embodiment, the execution subject can query the database corresponding to the business subject based on the row record information of the target data table and the attributes of each dimension, determine the attribute value of each dimension, and generate a target data table including at least the attribute value of each dimension. The target data table can also generate a target data table including the attribute value of each dimension by using the row record information of the target data table as the primary key and the attributes of each dimension as each column.
[0056] Continue to see Figure 3 , Figure 3It is a schematic diagram 300 of an application scenario of the data table generation method according to the present embodiment. The data table generation method of the present embodiment runs in an electronic device 301. First, the electronic device 301 determines at least one business data table associated with the business subject and row record information 302 of the target data table according to the business subject, and then the electronic device 301 counts the occurrence frequency of each field in each business data table, obtains at least one target field whose statistical value meets the preset condition, and uses each target field as each dimension of the target data table to generate each dimension table 303, then the electronic device 301 selects the dimension attribute 304 in each dimension table, and finally the electronic device 301 determines the attribute value of each dimension based on the row record information of the target data table and the attributes of each dimension, and generates a target data table 305 including at least the attribute value of each dimension.
[0057] The data table generation method provided by the embodiment of the present disclosure determines at least one business data table related to the business subject and row record information of the target data table according to the business subject, counts the occurrence frequency of each field in each business data table, obtains at least one target field whose statistical value meets the preset condition, and uses each target field as each dimension of the target data table to generate each dimension table, selects the dimension attribute in each dimension table, determines the attribute value of each dimension based on the row record information of the target data table and the attributes of each dimension, and generates a target data table including at least the attribute value of each dimension. A more comprehensive and effective data table generation method is realized.
[0058] Further references Figure 4 , which shows the process of another embodiment of the data table generation method. The process 400 of the data table generation method includes the following steps:
[0059] Step 401: Determine row record information of at least one business data table and a target data table associated with the business subject according to the business subject.
[0060] Step 402, using the word frequency-inverse document frequency technology to count the occurrence frequency of each field in the script library corresponding to each business data table, obtain at least one target field whose statistical value meets the preset conditions, and use each target field as each dimension of the target data table to generate each dimension table.
[0061] In this embodiment, the execution entity uses the term frequency-inverse document frequency technology (i.e., the TF-IDF algorithm technology) to count the occurrence frequencies of each field in the script library corresponding to each service data table, obtains at least one target field whose statistical value meets the preset conditions, and uses each target field as each dimension of the target data table to generate each dimension table. The statistical value represents the product of the term frequency TF value of the field and the inverse document frequency IDF value of the corresponding field. The preset conditions can be that the statistical value is greater than the threshold and / or the statistical value is before the preset serial number after all statistical values are sorted. The script library can include a historical query script library, a real-time task script library, etc.
[0062] Specifically, the statistical process of the target field includes: dividing each field in the script library according to each preset condition in multiple preset conditions to obtain various types of fields corresponding to each preset condition; using the term frequency-inverse document frequency technology to count the occurrence frequencies of each type of field, and obtaining at least one sub-target field of each type of field whose statistical value meets the corresponding preset condition; and combining the sub-target fields of various types of fields to obtain each target field.
[0063] In some alternative implementation manners of this embodiment, each preset condition is set based on each type of field information; each preset condition corresponds to one type of field information. The types of field information can include: basic information type, logistics information type, person information type, and other four types. By setting different preset conditions according to the type of field information, the high and low frequency standards for different categories are different. For example, for basic information type fields, because the occurrence frequency is relatively high, a higher frequency screening standard is set for basic information type fields, so as to realize the statistics of target fields with strong pertinence.
[0064] In some alternative implementation manners of this embodiment, before using the term frequency-inverse document frequency technology to count the occurrence frequencies of each field in each service data table in the script library and obtaining at least one target field whose statistical value meets the preset conditions, it further includes: performing stop word annotation on each field in the script library, so as to ignore each field marked as a stop word in the subsequent statistics of each field in the script library. "Stop words" in Chinese text mining refer to the words with the most occurrences, which are words that are of no help in finding the results and must be filtered out, such as the most commonly used words like "de" (的), "shi" (是), "zai" (在), etc. Here, query keywords such as SELECT, WHERE, FROM in the query script can be set as stop words, and script language keywords such as new, class, public, void, return in the real-time task script library can be set as stop words, so as to ignore each field marked as a stop word in the subsequent statistics of each field in the script library.
[0065] Step 403, select the dimension attributes in each dimension table.
[0066] In some optional implementations of this embodiment, before selecting the dimension attributes in each dimension table, it also includes: determining the associated fields related to each dimension according to the dimensions of the target data table; using the associated fields as the dimensions of the target data table, and merging them with the existing dimensions of the target data table to generate the dimensions of the final target data table. Further optimization is performed on the basis of the dimension construction of the original data table, such as associating the dimensions such as the star rating of the buyer and seller, the name of the store, and the category level into the fact table, making the fact table more complete and comprehensive, and improving the efficiency of filtering, querying, and statistical aggregation of the fact table.
[0067] Step 404: determine the attribute value of each dimension based on the row record information of the target data table and the attribute of each dimension, and generate a target data table including at least the attribute value of each dimension.
[0068] In some optional implementations of this embodiment, determining the row record information of the target data table includes: analyzing the business data table and selecting the information with the finest level of detail representing the business subject as the row record information of the target data table. By selecting the finest level of granularity, it is ensured that the application of the table has greater flexibility.
[0069] In some optional implementations of this embodiment, the method further includes: displaying the target data table for further confirmation and application of the data table generation result.
[0070] It should be noted that the above TF-IDF algorithm is a well-known technology that is currently widely studied and applied, and will not be described in detail here.
[0071] In this embodiment, the specific operations of steps 401, 403 and 404 are the same as Figure 2 The operations of steps 201, 203 and 204 in the illustrated embodiment are substantially the same and will not be described in detail herein.
[0072] from Figure 4 It can be seen that Figure 2Compared with the corresponding embodiment, the data table generation process 400 in this embodiment uses the word frequency-inverse document frequency technology to count the occurrence frequency of each field in the script library corresponding to each business data table, obtains at least one target field whose statistical value meets the preset conditions, and uses each target field as each dimension of the target data table to generate each dimension table. It avoids the field redundancy in the data table dimension confirmation in the prior art, saves computing resources, and solves the problem of high cost and incomplete field collection caused by collecting customer needs. By using the TF-IDF algorithm to identify the high and low frequency of use of each field in the data source, the model field is automatically selected, the model AP side coverage is improved, and the cost of collecting customer needs and the business experience threshold of R&D personnel are reduced. By setting different preset conditions, various high-frequency fields are used as the dimensions of the table, and the model is covered as much as possible. The construction of the data model is more flexible and targeted.
[0073] The terms used in the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. The singular forms of "a", "said" and "the" used in the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in this article refers to and includes any or all possible combinations of one or more associated listed items. It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0074] Further references Figure 5 , as a response to the above Figure 2 to Figure 4 The present disclosure provides an embodiment of a data table generating device, and the device embodiment is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0075] like Figure 5As shown, the data table generating device 500 of this embodiment includes: a first determining unit 501, a counting unit 502, a selecting unit 503 and a first generating unit 504. The first determining unit is configured to determine at least one business data table associated with the business subject and row record information of the target data table according to the business subject; the counting unit is configured to count the occurrence frequency of each field in each business data table, obtain at least one target field whose statistical value meets the preset condition, and use each target field as a dimension of the target data table to generate each dimension table; the selecting unit is configured to select dimension attributes in each dimension table; the first generating unit is configured to determine the attribute value of each dimension based on the row record information of the target data table and the attributes of each dimension, and generate a target data table including at least the attribute value of each dimension.
[0076] In this embodiment, the specific processing of the first determination unit 501, the statistical unit 502, the selection unit 503 and the first generation unit 504 of the data table generation device 500 and the technical effects thereof can be referred to in Figure 2 The relevant descriptions of steps 201 to 204 in the corresponding embodiment are not repeated here.
[0077] In some optional implementations of this embodiment, the statistical unit is further configured to use the word frequency-inverse document frequency technology to count the occurrence frequency of each field in each business data table in the script library corresponding to each business data table, and obtain at least one target field whose statistical value meets the preset conditions, and the statistical value represents the product of the word frequency of the field and the inverse document frequency of the corresponding field.
[0078] In some optional implementations of this embodiment, the device further includes: a marking unit configured to mark each field in the script library with stop words, so that the fields marked as stop words are ignored in subsequent statistics of each field in the script library.
[0079] In some optional implementations of this embodiment, the preset conditions in the statistical unit are that the statistical value is greater than a threshold value and / or the statistical value is located before a preset serial number after all statistical values are sorted; the statistical unit includes: a field division module, configured to divide the fields in the script library according to each preset condition in a plurality of preset conditions, and obtain various types of fields corresponding to each preset condition; a frequency statistics module, configured to use the word frequency-inverse document frequency technology to count the occurrence frequency of each type of field, and obtain at least one sub-target field of each type of field whose statistical value meets the corresponding preset condition; a merging module, configured to merge the sub-target fields of various types of fields to obtain each target field.
[0080] In some optional implementations of this embodiment, each preset condition in the statistical unit is set based on each type of field information; each preset condition in the statistical unit corresponds to a type of field information.
[0081] In some optional implementations of this embodiment, the device also includes: a second determination unit, configured to determine the associated fields related to each dimension based on the dimensions of the target data table; a second generation unit, configured to use the associated fields as the dimensions of the target data table, and merge them with the dimensions of the existing target data table to generate the dimensions of the final target data table.
[0082] In some optional implementations of this embodiment, the first determining unit is further configured to analyze the business data table and select information representing the business subject with the finest level of detail as row record information of the target data table.
[0083] In some optional implementations of this embodiment, the device further includes: a setting unit configured to use the row record information of the target data table as the primary key of the target data table.
[0084] In some optional implementations of this embodiment, the device further includes: a display unit configured to display the target data table.
[0085] Reference below Figure 6 , which shows a schematic diagram of the structure of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0086] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0087] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0088] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0089] It should be noted that the computer-readable medium of the embodiment of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiment of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the embodiment of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0090] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: determines at least one business data table associated with the business subject and row record information of the target data table according to the business subject; counts the occurrence frequency of each field in each business data table, obtains at least one target field whose statistical value meets the preset conditions, and uses each target field as each dimension of the target data table to generate each dimension table; selects dimension attributes in each dimension table; determines the attribute value of each dimension based on the row record information of the target data table and the attributes of each dimension, and generates a target data table including at least the attribute value of each dimension.
[0091] Computer program code for performing the operations of embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as a separate software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0092] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0093] The units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The units described may also be provided in a processor, for example, may be described as: a processor comprising a first determination unit, a statistics unit, a selection unit, and a first generation unit. The names of these units do not constitute limitations on the units themselves in certain circumstances, for example, the first determination unit may also be described as "a unit for determining at least one business data table associated with the business subject and row record information of a target data table according to the business subject".
[0094] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.
Claims
1. A method for generating a data table, comprising: Determine, according to the business subject, at least one business data table and row record information of a target data table associated with the business subject; Counting the occurrence frequency of each field in each of the business data tables to obtain at least one target field whose statistical value meets a preset condition, and using each of the target fields as each dimension of the target data table to generate each dimension table; Select dimension attributes in each dimension table; Based on the row record information of the target data table and the attributes of each dimension, the attribute value of each dimension is determined, and the target data table at least including the attribute value of each dimension is generated.
2. The method according to claim 1, wherein: The method of counting the occurrence frequency of each field in each of the business data tables to obtain at least one target field whose statistical value satisfies a preset condition includes: The word frequency-inverse document frequency technology is used to count the occurrence frequency of each field in the script library corresponding to each business data table, and at least one target field whose statistical value meets the preset conditions is obtained. The statistical value represents the product of the word frequency of the field and the inverse document frequency of the corresponding field.
3. The method according to claim 2, wherein: The preset condition is that the statistical value is greater than a threshold value and / or the statistical value is located before a preset sequence number after all statistical values are sorted; the word frequency-inverse document frequency technology is used to count the occurrence frequency of each field in the script library corresponding to each business data table to obtain at least one target field whose statistical value meets the preset condition, including: According to each preset condition in the plurality of preset conditions, each field in the script library is divided to obtain each type of field corresponding to each preset condition; The frequency of occurrence of each type of field is counted using the word frequency-inverse document frequency technology to obtain at least one sub-target field of each type of field whose statistical value meets the corresponding preset conditions; The sub-target fields of the various fields are merged to obtain the target fields.
4. The method according to claim 3, wherein: The preset conditions are set based on the types of field information; each preset condition corresponds to one type of field information.
5. The method according to claim 1, wherein: Before selecting the dimension attributes in each dimension table, the method further includes: According to each dimension of the target data table, determining associated fields related to each dimension; The associated fields are used as dimensions of the target data table, and are merged with the existing dimensions of the target data table to generate the final dimensions of the target data table.
6. A data table generating device, comprising: A first determining unit is configured to determine row record information of at least one business data table and a target data table associated with the business subject according to the business subject; A statistical unit is configured to count the occurrence frequency of each field in each of the business data tables, obtain at least one target field whose statistical value meets a preset condition, and use each of the target fields as each dimension of the target data table to generate each dimension table; A selection unit is configured to select dimension attributes in each dimension table; The first generating unit is configured to determine the attribute value of each dimension based on the row record information of the target data table and the attribute of each dimension, and generate the target data table at least including the attribute value of each dimension.
7. The device according to claim 6, wherein: The statistical unit is further configured to use the word frequency-inverse document frequency technology to count the occurrence frequency of each field in the script library corresponding to each business data table, and obtain at least one target field whose statistical value meets the preset conditions, and the statistical value represents the product of the word frequency of the field and the inverse document frequency of the corresponding field.
8. The device according to claim 7, wherein: The preset condition in the statistical unit is that the statistical value is greater than a threshold value and / or the statistical value is located before a preset sequence number after all statistical values are sorted; The statistical unit comprises: A field division module is configured to divide each field in the script library according to each preset condition in a plurality of preset conditions to obtain various types of fields corresponding to each of the preset conditions; A frequency statistics module is configured to use a word frequency-inverse document frequency technique to count the occurrence frequency of each type of field, and obtain at least one sub-target field of each type of field whose statistical value meets corresponding preset conditions; The merging module is configured to merge the sub-target fields of the various types of fields to obtain the target fields.
9. The device according to claim 8, wherein: The preset conditions in the statistical unit are set based on the types of field information; each preset condition in the statistical unit corresponds to a type of field information.
10. The apparatus according to claim 6, further comprising: A second determining unit is configured to determine, according to each dimension of the target data table, an associated field related to each dimension; The second generating unit is configured to use the associated fields as dimensions of the target data table, and merge them with the existing dimensions of the target data table to generate the final dimensions of the target data table.
11. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
12. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method for rapidly collecting multi-layer fact data based on SQL statements
CN104021156A
Data import method and device, computer equipment and storage medium
CN112685415A