A data library splitting method, system, electronic device and computer program product
By constructing a mapped directed acyclic graph and monitoring state stability, the problems of scattered storage of project data and high query latency caused by unreasonable data sharding were solved, thereby improving the reliability and query efficiency of data sharding.
Patent Information
- Application Number
- CN202511421227.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-30
AI Technical Summary
In existing technologies, unreasonable data partitioning leads to scattered storage of project data and high query latency.
By acquiring project information from the target company, spatial reasoning is performed through the long-range spatial reasoning channel in the dual-architecture data sharding platform. A mapped directed acyclic graph is constructed, and state stability is monitored over time to determine the data sharding scheme and build a mapping routing table.
It improves the reliability of data sharding and query reliability, reduces query latency, and optimizes data storage and query efficiency.
Smart Images

Figure CN120910161B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a data library division method, system, electronic device and computer program product. BACKGROUND
[0002] At present, the data library division method usually adopts a horizontal library division strategy based on time or regional hash. These methods divide data into different physical libraries to improve the scalability and performance of the system. However, the traditional horizontal library division strategy has some deficiencies when facing complex project data. For example, the hash strategy based on time or region cannot ensure the integrity of the data of a single project in the same physical library, which may lead to frequent cross-shard transactions. This not only increases the complexity of the system, but also may significantly increase the query delay, affecting the overall data processing efficiency.
[0003] There is a technical problem in the prior art that due to unreasonable data library division, project data is stored in a scattered manner, and the query delay is high. SUMMARY
[0004] The present application provides a data library division method, system, electronic device and computer program product, which solves the technical problem in the prior art that due to unreasonable data library division, project data is stored in a scattered manner, and the query delay is high.
[0005] The first aspect of the present application provides a data library division method, which comprises:
[0006] Obtain K project information of a target enterprise, call a long-range spatial reasoning channel in a data library division double-architecture platform to perform spatial reasoning on the K project information, obtain a K project information interaction primitive set and a K project information interaction primitive association direction set, wherein K is a positive integer; take each project information interaction primitive in the K project information interaction primitive set as a node, determine a directed edge according to the K project information interaction primitive association direction, and construct a mapping directed acyclic graph; call a query response processing channel in the data library division double-architecture platform to perform state stability time sequence iteration monitoring on each node in the mapping directed acyclic graph, and obtain a plurality of node state stability identification results; identify each node of the mapping directed acyclic graph according to the plurality of node state stability identification results, and obtain an identified mapping directed acyclic graph; perform library division scheme identification based on the identified mapping directed acyclic graph, determine a data library division scheme, and construct and store a mapping routing table based on the data library division scheme.
[0007] In a possible implementation, K project information of a target enterprise is acquired, a long-range spatial reasoning channel in a data split-database dual architecture platform is called to perform spatial reasoning on the K project information, and K project information interaction element sets and K project information interaction element association direction sets are obtained, including: performing multi-modal data extraction on the K project information by using a data extraction branch of the long-range spatial reasoning channel, and obtaining K multi-modal data sets; generating data universal descriptions on the K multi-modal data sets by using a data description branch of the long-range spatial reasoning channel, and obtaining K data universal description text sets, where each data universal description text includes an object, an attribute, a relationship, and a scene; performing fine-grained spatial reasoning question answering on the K data universal description text sets, and extracting the K project information interaction element sets and the K project information interaction element association direction sets.
[0008] In a possible implementation, the fine-grained spatial reasoning question answering on the K data universal description text sets, and the extraction of the K project information interaction element sets and the K project information interaction element association direction sets, include: performing keyword extraction on the K data universal description text sets respectively, and obtaining K keyword vector sets; performing problem dispersion retrieval in a fine-grained spatial reasoning question answering library based on the K keyword vector sets, and determining K fine-grained spatial reasoning question group sets; performing reply analysis on the K fine-grained spatial reasoning question group sets based on the K data universal description text sets as reply materials, and obtaining K data universal description question reply reliability sets; performing screening on the K data universal description text sets based on the K data universal description question reply reliability sets, and determining K effective data universal description text sets; performing interaction element and interaction element association direction extraction on the K effective data universal description text sets by using spaCy, and obtaining the K project information interaction element sets and the K project information interaction element association direction sets.
[0009] In a possible implementation, the problem dispersion retrieval in the fine-grained spatial reasoning question answering library based on the K keyword vector sets, and the determination of the K fine-grained spatial reasoning question group sets, include: extracting a first keyword vector from the K keyword vector sets, and performing matching in the fine-grained spatial reasoning question answering library according to the first keyword vector, to obtain a first matched fine-grained spatial reasoning question group; performing problem dispersion retrieval on the first matched fine-grained spatial reasoning question group, to determine a first fine-grained spatial reasoning question group; and adding the first fine-grained spatial reasoning question group into the K fine-grained spatial reasoning question group sets.
[0010] In a possible implementation, the problem dispersion search is performed on the first matched fine-grained spatial reasoning problem set, the first fine-grained spatial reasoning problem set is determined, including: randomly extracting M matched fine-grained spatial reasoning problems from the first matched fine-grained spatial reasoning problem set, where M is a positive integer; performing pairwise enumeration on the M matched fine-grained spatial reasoning problems and performing dispersion calculation on the enumeration combinations to obtain an enumeration combination dispersion set, and performing the same association matched fine-grained spatial reasoning problem division and mean value calculation on the enumeration combination dispersion set to obtain M matched fine-grained spatial reasoning problem dispersion means; performing problem dispersion verification based on the M matched fine-grained spatial reasoning problem dispersion means, and if the verification fails, obtaining an abnormal matched fine-grained spatial reasoning problem and a to-be-supplemented matched fine-grained spatial reasoning problem set; taking the abnormal matched fine-grained spatial reasoning problem as an index, performing dispersion calculation on the matched fine-grained spatial reasoning problems in the first matched fine-grained spatial reasoning problem set except for the M matched fine-grained spatial reasoning problems, and performing screening according to a preset dispersion threshold and ordering according to the dispersion from large to small to obtain a candidate matched fine-grained spatial reasoning problem sequence; sequentially adding the candidate matched fine-grained spatial reasoning problem sequence into the to-be-supplemented matched fine-grained spatial reasoning problem set, and performing problem dispersion verification until the verification passes, to obtain the first fine-grained spatial reasoning problem set.
[0011] In a possible implementation, the state stability time sequence iteration monitoring is performed on each node in the mapping directed acyclic graph by calling the query response processing channel of the data split database dual architecture platform, a plurality of node state stability identification results are obtained, including: performing state monitoring on each node in the mapping directed acyclic graph within a preset monitoring window according to a preset query response index through the query response processing channel to obtain a plurality of node state feature sequences; performing in-sequence adjacent feature iteration interaction on the plurality of node state feature sequences to obtain a plurality of interaction node state features; performing state stability identification based on the plurality of interaction node state features to obtain a plurality of node state stability identification results.
[0012] In a possible implementation, the plurality of node state feature sequences are subjected to intra-sequence adjacent feature iterative interaction to obtain a plurality of interactive node state features, including: respectively performing adjacent sub-feature mapping deviation identification on a plurality of first-position node state features and a plurality of second-position node state features in the plurality of node state feature sequences to obtain a plurality of first adjacent sub-feature deviation sets; performing normalization processing on the plurality of first adjacent sub-feature deviation sets and filling the adjacent sub-feature deviation sets into an initially empty adjacent matrix to obtain a plurality of first adjacent matrices; respectively performing iterative interaction on a plurality of second-position node state features based on the plurality of first adjacent matrices to obtain a plurality of first iterative interactive node state features, and performing adjacent feature iterative interaction on a plurality of third-position node state features in the plurality of node state feature sequences by using the plurality of first iterative interactive node state features, and iteratively obtaining the plurality of interactive node state features in the same manner.
[0013] In a second aspect of the present disclosure, a data library splitting system is provided, and the system includes:
[0014] The spatial reasoning module is configured to obtain K pieces of project information of a target enterprise, call a long-range spatial reasoning channel in the data library splitting dual-architecture platform to perform spatial reasoning on the K pieces of project information, and obtain a K-piece project information interaction primitive set and a K-piece project information interaction primitive association direction set, where K is a positive integer. The mapping directed acyclic graph construction module is configured to take each project information interaction primitive in the K-piece project information interaction primitive set as a node, determine a directed edge according to the K-piece project information interaction primitive association direction, and construct a mapping directed acyclic graph. The stability identification result obtaining module is configured to call a query response processing channel in the data library splitting dual-architecture platform to perform state stability time sequence iterative monitoring on each node in the mapping directed acyclic graph, and obtain a plurality of node state stability identification results. The identification module is configured to identify each node of the mapping directed acyclic graph according to the plurality of node state stability identification results, and obtain an identified mapping directed acyclic graph. The mapping routing table construction and storage module is configured to identify a data library splitting scheme based on the identified mapping directed acyclic graph, determine the data library splitting scheme, and construct and store a mapping routing table based on the data library splitting scheme.
[0015] In a third aspect of the present disclosure, an electronic device is provided, including a memory and a processor, the memory stores executable instructions, and the processor executes the executable instructions stored in the memory to implement any step of the first aspect of the present disclosure.
[0016] In a fourth aspect of the present disclosure, a computer program product is provided, which includes computer instructions, and the computer instructions are executed by a processor to implement any step of the first aspect of the present disclosure.
[0017] The one or more technical solutions provided in the application have at least the following technical effects or advantages:
[0018] The application obtains K project information of a target enterprise, calls a long-range space reasoning channel in a data split-database dual-architecture platform to perform space reasoning on the K project information, obtains a K project information interaction primitive set and a K project information interaction primitive association direction set, wherein K is a positive integer; takes each project information interaction primitive in the K project information interaction primitive set as a node, determines a directed edge according to the K project information interaction primitive association direction, and constructs a mapping directed acyclic graph; calls a query response processing channel in the data split-database dual-architecture platform to perform state stability time sequence iteration monitoring on each node in the mapping directed acyclic graph, and obtains a plurality of node state stability identification results; identifies each node of the mapping directed acyclic graph according to the obtained plurality of node state stability identification results, and obtains an identified mapping directed acyclic graph; identifies a split-database scheme based on the identified mapping directed acyclic graph, determines a data split-database scheme, and constructs and stores a mapping routing table based on the data split-database scheme. The technical effect of improving data split-database reliability and query reliability by determining the interaction primitives and association directions of project information is achieved.
[0019] The above description is only a summary of the technical solutions of the application. In order to more clearly understand the technical means of the application, the application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A data split-database method flowchart is provided for the embodiments of the application.
[0021] Figure 2 A data split-database system structure diagram is provided for the embodiments of the application.
[0022] Figure 3 An internal structure diagram of an electronic device is provided for the embodiments of the application.
[0023] BRIEF DESCRIPTION OF DRAWINGS
[0024] Space reasoning module 11, mapping directed acyclic graph construction module 12, stability identification result obtaining module 13, identification module 14, mapping routing table construction and storage module 15, bus 300, receiver 301, processor 302, transmitter 303, memory 304, bus interface 305. DETAILED DESCRIPTION
[0025] The embodiment of the application provides a data library division method, and solves the technical problem of high query delay caused by dispersed storage of project data due to unreasonable data library division in the prior art.
[0026] After introducing the basic principle of the application, various non-limiting embodiments of the application will be specifically introduced in combination with the accompanying drawings of the specification. It should be understood that the specific embodiments described herein are only used to explain the application and not to limit the application.
[0027] Embodiment one, as shown in the figure, the embodiment of the application provides a data library division method, the method comprises: Figure 1
[0028] Step S100: obtaining K project information of a target enterprise, calling a long-range space reasoning channel in a data library double-architecture platform to perform space reasoning on the K project information, obtaining a K project information interaction primitive set and a K project information interaction primitive association direction set, wherein K is a positive integer;
[0029] Further, obtaining K project information of a target enterprise, calling a long-range space reasoning channel in a data library double-architecture platform to perform space reasoning on the K project information, obtaining a K project information interaction primitive set and a K project information interaction primitive association direction set, the step S100 of the embodiment of the application further comprises:
[0030] using a data extraction branch of the long-range space reasoning channel to perform multi-modal data extraction on the K project information, obtaining K multi-modal data sets;
[0031] using a data description branch of the long-range space reasoning channel to generate data universal descriptions for the K multi-modal data sets, obtaining a K data universal description text set, wherein each data universal description text comprises an object, an attribute, a relationship and a scene;
[0032] iterating the K data universal description text set to perform fine-grained space reasoning question answering, extracting a K project information interaction primitive set and a K project information interaction primitive association direction set.
[0033] It should be noted that the data split database double-architecture platform is a system architecture supporting data split database processing, including two independent system modules, which are a long-range spatial reasoning channel and a query response processing channel. Among them, the long-range spatial reasoning module is used for high-level data reasoning and abstraction of project information, and the query response processing channel focuses on fast response and database operation. The target enterprise is any enterprise that needs to process data in a split database, such as a construction enterprise, an information-based enterprise, etc. The K project information is the data corresponding to the projects being carried out or having been completed by the target enterprise, which includes various types of data, such as structured, semi-structured and unstructured data.
[0034] The long-range spatial reasoning channel in the data split database double-architecture platform is used to perform spatial reasoning on the K project information respectively, and the scattered and multi-type project raw information is refined and reasoned, thereby reducing redundant data and improving the accuracy of data split database. Since there may be cross-terms between the K project information, they often appear simultaneously during data query, so not only the key information in each project information needs to be captured, but also the relevance between different data needs to be considered when performing data split database. After completing the spatial reasoning, the K project information can obtain a K project information interaction primitive set and a K project information interaction primitive association direction set. Among them, the K project information interaction primitive set refers to the basic unit that constitutes the project data association, such as project subject, associated unit, core attribute, etc., and the K project information interaction primitive association direction set refers to the logical pointing relationship between interaction primitives, such as: project → supplier, indicating the dependency direction of the project to the supplier.
[0035] Preferably, the long-range spatial reasoning channel includes a data extraction branch and a data description branch, which are obtained after further training based on a framework constructed based on a convolutional neural network. Among them, the data extraction branch is used to extract multi-modal data in the K project information respectively to obtain the K multi-modal data set. The data description branch is used to describe the K multi-modal data set in text. Because compared with the complex and diversified data expression form in multi-modal data, the text natural language has stronger semantic bearing capacity and expression concentration, so by describing the object, attribute, relationship and scene of each multi-modal data in text, the K data general description text set is obtained, which achieves the technical effect of facilitating subsequent in-depth capture of key data in project information.
[0036] In one embodiment, a plurality of sample item information and a plurality of sample multi-modal data are obtained, and a plurality of sample data general description texts in a general format, i.e., an object, an attribute, a relationship, and a scene description format, are textually described by a person skilled in the art. The framework based on the convolutional neural network is supervised and trained by using the plurality of sample item information and the plurality of sample multi-modal data. In the training process, the framework parameters are updated and adjusted according to the accuracy of the output until convergence is achieved, and the trained data extraction branch is obtained. Based on the same principle, the data description branch is trained and constructed based on the plurality of sample multi-modal data and the plurality of sample data general description text set.
[0037] By performing multi-modal data extraction, description, and reasoning, the K project information of the target enterprise is comprehensively understood, and the key components of the project information are identified, which lays a foundation for optimizing data storage and reducing query delay.
[0038] Further, the K data general description text set is traversed to perform fine-grained spatial reasoning question and answer, and the K project information interaction primitive set and the K project information interaction primitive association direction set are extracted. The embodiment of the present application further includes the following steps S100:
[0039] The K data general description text set is respectively used for keyword extraction to obtain a K keyword vector set;
[0040] Based on the K keyword vector set, question dispersion retrieval is performed in the fine-grained spatial reasoning question and answer library to determine a K fine-grained spatial reasoning question group set;
[0041] The K data general description text set is used as a reply material to reply and analyze the K fine-grained spatial reasoning question group set to obtain a K data general description question and reply reliability set;
[0042] Based on the K data general description question and reply reliability set, the K data general description text set is screened to determine a K effective data general description text set;
[0043] The K effective data general description text set is extracted by using spaCy to obtain a K project information interaction primitive set and a K project information interaction primitive association direction set.
[0044] In one embodiment, keyword extraction is performed on the K sets of data general description text using Word2Vec to obtain K sets of keyword vectors. Semantic quantization is performed to lay the foundation for subsequent question dispersion retrieval in the fine-grained space reasoning question and answer library. For example, a project management company has K=3 commissioned projects, project X: A group IT system integration management, project Y: B mall activity planning and execution, and project Z: C factory workshop renovation supervision. First, keyword extraction is performed on the data general description text of the three projects. For example, the description text of project X is extracted as "IT system integration management, budget 4 million, construction period 100 days, customer A group, subcontractor B technical team, demand research stage, acceptance standard includes data migration test". Each keyword is converted into a 128-dimensional numerical vector using the BERT word embedding model to form three sets of keyword vectors. The fine-grained space reasoning question and answer library is a database that contains fine-grained reasoning questions and corresponding answer example structures and corresponding sample description text keyword vectors. The sample description text keyword vectors are matched according to the input keyword vectors, and the matching results are subjected to question dispersion retrieval to determine the corresponding fine-grained space reasoning question group.
[0045] Since the data general description text is the explicit information of the project information, in order to mine the implicit information of the project information, fine-grained space reasoning questions can be used to mine the deep associations between different projects or different nodes of the same project based on key information such as objects, attributes, actions, locations, and interactions in the data general description text. The data general description text is used as a reference to answer the corresponding fine-grained space reasoning question group using a deep learning model. The questions that can be clearly determined are answered and analyzed to obtain K sets of data general description question answer reliability.
[0046] Preferably, a plurality of sample fine-grained space reasoning questions and corresponding sample data general description texts are obtained, and a plurality of sample answer results are obtained by a person skilled in the art based on the plurality of sample data general description texts to answer the plurality of sample fine-grained space reasoning questions, thereby obtaining model training data. The deep learning model is supervised trained based on the model training data to obtain a trained deep learning model. The deep learning model is used to answer the K sets of fine-grained space reasoning question groups based on the K sets of data general description texts to obtain a plurality of answer results. The answer results are compared according to the corresponding answer example structure in the fine-grained space reasoning question and answer library, and the corresponding data general description question answer reliability is obtained according to the structure comparison success rate.
[0047] The pre-set data general description question answer reliability threshold value preset by the person skilled in the art is obtained, the data general description text corresponding to the data general description question answer reliability greater than or equal to the pre-set data general description question answer reliability threshold value in the K data general description question answer reliability set is added to the K effective data general description text set. The K data general description texts are cut into single words, specifically, the part-of-speech of each word in the text is tagged, and the grammatical role of each word is identified. Then, syntactic analysis is performed to identify the relationship between the words and determine the grammatical structure of each part of the sentence. Further, spaCy is used for named entity recognition to extract key interaction primitives from the text. These interaction primitives can be entity names such as project modules: servers, offices, conference rooms, etc.; components: storage units, network routers, etc.; task nodes: computing nodes, storage nodes, etc. Thus, the core elements involved in the project description are identified, and the K project information interaction primitive set is obtained.
[0048] Further, through dependency syntax analysis, the relationship between each word in the text is mined. For example, a project includes multiple offices, the project is the subject, the inclusion is the verb, and the multiple offices are the object. The office can be extracted to depend on the project, and the computing node depends on the computing task, so as to effectively determine the dependency relationship and interaction path between the interaction primitives, that is, to determine the association direction between the interaction primitives. Therefore, through dependency syntax analysis, the K project information interaction primitive set and the K project information interaction primitive association direction set can be extracted from the effective data general description text. Among them, the project information interaction primitive association direction describes the dependency or data flow direction between different interaction primitives, for example, storage unit←data transmission, indicating that the storage unit is responsible for processing data transmission, and the data transmission depends on the storage unit. By identifying the interaction primitives and the interaction primitive association direction, the technical effect of providing data support for subsequent construction of a mapping directed acyclic graph is achieved.
[0049] Further, based on the K key word vector set, question dispersion retrieval is performed in the fine-grained space reasoning question and answer library to determine the K fine-grained space reasoning question group set, and the embodiment of the present application further includes the following steps.
[0050] A first key word vector is extracted from the K key word vector set, and matching is performed in the fine-grained space reasoning question and answer library according to the first key word vector to obtain a first matched fine-grained space reasoning question group.
[0051] The first matched fine-grained space reasoning question group is subjected to question dispersion retrieval to determine a first fine-grained space reasoning question group.
[0052] The first fine-grained space reasoning question group is added to the K fine-grained space reasoning question group set.
[0053] Further, the first matching fine-grained spatial reasoning problem group is subjected to problem dispersion search to determine a first fine-grained spatial reasoning problem group. The step S100 of the embodiment of the present application further includes:
[0054] M matching fine-grained spatial reasoning problems are randomly extracted from the first matching fine-grained spatial reasoning problem group, where M is a positive integer.
[0055] The M matching fine-grained spatial reasoning problems are subjected to pairwise enumeration and dispersion calculation of the enumerated combinations to obtain an enumerated combination dispersion set. The enumerated combination dispersion set is subjected to matching fine-grained spatial reasoning problem division and mean calculation of the same association to obtain M matching fine-grained spatial reasoning problem dispersion means.
[0056] The M matching fine-grained spatial reasoning problem dispersion means are subjected to problem dispersion verification. If the verification fails, an abnormal matching fine-grained spatial reasoning problem and a to-be-supplemented matching fine-grained spatial reasoning problem group are obtained.
[0057] The abnormal matching fine-grained spatial reasoning problem is used as an index to calculate the dispersion of the matching fine-grained spatial reasoning problems in the first matching fine-grained spatial reasoning problem group except for the M matching fine-grained spatial reasoning problems. The dispersion is screened according to a preset dispersion threshold and sorted in descending order of dispersion to obtain a candidate matching fine-grained spatial reasoning problem sequence.
[0058] The candidate matching fine-grained spatial reasoning problem sequence is sequentially added to the to-be-supplemented matching fine-grained spatial reasoning problem group, and subjected to problem dispersion verification until the verification passes to obtain a first fine-grained spatial reasoning problem group.
[0059] In the embodiment of the present application, the first in the first keyword vector does not represent the order, but refers to any one of the K keyword vector sets. The cosine similarity formula is used to calculate the similarity between the first keyword vector and the sample description text keyword vector in the fine-grained spatial reasoning question and answer library. The fine-grained reasoning problem corresponding to the sample description text keyword vector that meets the similarity threshold set by the person skilled in the art in the calculation result is added to the first fine-grained reasoning problem group.
[0060] Further, the first fine-grained reasoning problem group is subjected to problem dispersion search, that is, the repeated and approximate fine-grained reasoning problems in the first fine-grained reasoning problem group are removed, and the more dispersed fine-grained reasoning problems are selected, thereby ensuring the diversity of the questions and achieving the technical effect of improving the depth of mining of the data general description text.
[0061] The M matching fine-grained spatial reasoning problems are randomly extracted from the first matching fine-grained spatial reasoning problem group in a random selection manner, where M is a positive integer. Then, in order to compare the dispersion degrees between different matching fine-grained spatial reasoning problems in the M matching fine-grained spatial reasoning problems, enumeration combinations are obtained through two-by-two enumeration. The similarity of each enumeration combination in the enumeration combination set is calculated by using the cosine similarity formula, and the difference between the calculation result and 1 is taken as the enumeration combination dispersion degree, thereby obtaining an enumeration combination dispersion degree set. Each enumeration combination dispersion degree reflects the dispersion degree of each enumeration combination. Furthermore, the M matching fine-grained spatial reasoning problems are taken as indexes, and the enumeration combination dispersion degrees associated with the matching fine-grained spatial reasoning problems in the enumeration combination dispersion degree set are added to the corresponding set, thereby obtaining M associated enumeration combination dispersion degree sets. Then, the mean of the M associated enumeration combination dispersion degree sets is calculated to obtain the dispersion degree mean of the M matching fine-grained spatial reasoning problems. The dispersion degree mean of the M matching fine-grained spatial reasoning problems reflects the representativeness of the M matching fine-grained spatial reasoning problems. The higher the dispersion degree, the lower the repetition of the corresponding matching fine-grained spatial reasoning problem with other matching fine-grained spatial reasoning problems, and the more representative it is.
[0062] It is verified whether there is a matching fine-grained spatial reasoning problem dispersion degree mean less than the preset dispersion degree threshold set by the person skilled in the art in the M matching fine-grained spatial reasoning problem dispersion degree mean. If yes, the verification fails, and there are similar M matching fine-grained spatial reasoning problems in the obtained M matching fine-grained spatial reasoning problems. At this time, the matching fine-grained spatial reasoning problem corresponding to the minimum value in the M matching fine-grained spatial reasoning problem dispersion degree mean is taken as an abnormal matching fine-grained spatial reasoning problem, and the remaining M-1 matching fine-grained spatial reasoning problems are added to the to-be-supplemented matching fine-grained spatial reasoning problem group.
[0063] Furthermore, based on the same principle as obtaining the enumeration combination dispersion degree, the dispersion degrees of the matching fine-grained spatial reasoning problems in the first matching fine-grained spatial reasoning problem group except the M matching fine-grained spatial reasoning problems are calculated, and the dispersion degrees of the matching fine-grained spatial reasoning problems are identified with the abnormal matching fine-grained spatial reasoning problem. The matching fine-grained spatial reasoning problems with a dispersion degree less than the preset dispersion degree threshold are removed from the identification result, and the remaining matching fine-grained spatial reasoning problems are sorted in descending order of dispersion degree to obtain a candidate matching fine-grained spatial reasoning problem sequence.
[0064] Then, the candidate matching fine-grained space reasoning problem sequence is added into the to-be-supplemented matching fine-grained space reasoning problem group in sequence, and verification is performed based on the same principle as the above problem dispersion verification until the verification is passed, and a first fine-grained space reasoning problem group is obtained. Further, the first fine-grained space reasoning problem group is added into the K fine-grained space reasoning problem group set. By performing problem dispersion verification, the diversity of the fine-grained space reasoning problem group can be guaranteed, thereby effectively identifying deep-level implicit information of data, and also avoiding interference of redundant information.
[0065] Step S200: taking each item information interaction element in the K item information interaction element set as a node, determining a directed edge according to the K item information interaction element association direction, and constructing a mapping directed acyclic graph;
[0066] In one possible embodiment, the mapping directed acyclic graph is a graph structure constructed with the item information interaction element as a node and the item information interaction element association direction as a directed edge. Since there is no cyclic path from a node to the node again in the mapping directed acyclic graph, the business logic of stage progression and one-way dependence in project management can be met, such as one-way association of project→budget→cost accounting without cyclic dependence. For example, a project management enterprise has K=2 projects, which are project P: D company digital transformation project management and project Q: E park smart operation project management. The interaction element set of project P is {project P, D company, ERP supplier, budget 1200 million, data security compliance standard, demand research stage, system deployment stage}, and the interaction element set of project Q is {project Q, E park, security equipment supplier, budget 800 million, smart lighting standard, equipment installation stage, operation and maintenance training stage}. According to the association direction project P→D company, the directed edge relationship is the customer associated with demand docking, project P→ERP supplier, and the directed edge relationship is the collaboration associated with system supply, etc. Taking each item information interaction element in the K item information interaction element set as a node, determining a directed edge according to the K item information interaction element association direction, and when a plurality of directed edges are determined, integrating the nodes and the directed edges, checking and excluding the cyclic path, such as avoiding the invalid cycle of project P→budget 1200 million→project P, and ensuring that all directed edges meet the one-way logic of subject→object and predecessor→successor, to form an enterprise-level mapping directed acyclic graph.
[0067] By converting the dispersed project interaction elements and the project information interaction element association direction into a structured and visual directed acyclic graph, the hierarchical dependence and logical relationship of the internal data of the project are intuitively presented, the limitation of data isolated storage in the traditional database strategy is avoided, the technical effect of providing an association model basis for subsequent node state monitoring and database scheme identification is achieved, and a graphical analysis framework is established to solve the cross-shard transaction and query delay problem.
[0068] Step S300: calling a query response processing channel in the data split database dual architecture platform, performing state stability time sequence iteration monitoring on each node in the mapping directed acyclic graph, and obtaining a plurality of node state stability identification results;
[0069] Further, the query response processing channel in the data split database dual architecture platform is called to perform state stability time sequence iteration monitoring on each node in the mapping directed acyclic graph, and a plurality of node state stability identification results are obtained. The step S300 of the embodiment of the application further includes:
[0070] The state of each node in the mapping directed acyclic graph is monitored in a preset monitoring window according to a preset query response index through the query response processing channel, and a plurality of node state feature sequences are obtained.
[0071] The plurality of node state feature sequences are subjected to intra-sequence adjacent feature iteration interaction, and a plurality of interaction node state features are obtained.
[0072] Based on the plurality of interaction node state features, state stability identification is performed, and a plurality of node state stability identification results are obtained.
[0073] Further, the plurality of node state feature sequences are subjected to intra-sequence adjacent feature iteration interaction, and a plurality of interaction node state features are obtained. The step S300 of the embodiment of the application further includes:
[0074] The plurality of first node state features and the plurality of second node state features in the plurality of node state feature sequences are respectively subjected to adjacent sub-feature mapping deviation identification, and a plurality of first adjacent sub-feature deviation sets are obtained.
[0075] The plurality of first adjacent sub-feature deviation sets are subjected to normalization processing and filled into an initially empty adjacent matrix, and a plurality of first adjacent matrices are obtained.
[0076] Based on the plurality of first adjacent matrices, the plurality of second node state features are respectively subjected to iteration interaction, a plurality of first iteration interaction node state features are obtained, and the plurality of first iteration interaction node state features are used to perform adjacent feature iteration interaction on the plurality of third node state features in the plurality of node state feature sequences. In this way, the plurality of interaction node state features are obtained.
[0077] It should be noted that the query response processing channel is a module responsible for real-time query and transaction response, mainly processing query requests from users or systems and responding to query results. Through the query response processing channel, the state of each node in the mapping directed acyclic graph can be monitored within the preset monitoring window set by the person skilled in the art according to the preset query response indicators, so as to obtain a plurality of node state feature sequences that can reflect the query response of each node. Among them, the preset query response indicators include query delay, query data consistency and node load.
[0078] Adjacent sub-feature mapping deviation identification is performed on the plurality of first node state features and the plurality of second node state features in the plurality of node state feature sequences, that is, deviation identification is performed on the node state sub-features of the same type in the adjacent sub-features, such as query delay, and the deviation identification result is taken as a sub-feature deviation degree, and a plurality of first adjacent sub-feature deviation degree sets are obtained.
[0079] Further, the plurality of first adjacent sub-feature deviation degree sets are normalized by using the maximum-minimum normalization method to eliminate the dimension influence of different sub-features, and then filled into the initially empty matrix to obtain a plurality of first adjacent matrices. The plurality of adjacent matrices reflect the node state feature deviation between adjacent two nodes. The plurality of first iteration interaction node state features are obtained by using the graph convolutional neural network to iteratively interact the plurality of first adjacent matrices and the plurality of second node state features. The plurality of first iteration interaction node state features respectively implicitly reflect the deviation between the plurality of first node state features and the plurality of second node state features, that is, the fluctuation. Then, based on the same principle, the plurality of third node state features in the plurality of node state feature sequences are iteratively interacted by using the plurality of first iteration interaction node state features, and so on, to obtain the plurality of interaction node state features. The state stability identifier is constructed, preferably, a plurality of sample interaction node state features and a plurality of sample node state stability identification results are obtained as identifier training data, and the framework constructed based on the feedforward neural network is supervised trained by using the identifier training data until the training converges, and the trained state stability identifier is obtained.
[0080] The state stability identifier is called to identify the state stability of the plurality of interaction node state features, and a plurality of node state stability identification results are obtained. The node state stability identification result includes stable and unstable, and if it is stable, the possibility of inconsistency or performance bottleneck caused by the current storage path of the corresponding node is lower.
[0081] Step S400: Identify each node of the mapped directed acyclic graph based on the results of multiple node state stability identification to obtain the identified mapped directed acyclic graph;
[0082] Step S500: Identify the database sharding scheme based on the directed acyclic graph of the identifier mapping, determine the database sharding scheme, and construct and store the mapping routing table based on the database sharding scheme.
[0083] In one possible embodiment, the stability identification results of multiple nodes are respectively labeled to each node in a mapped directed acyclic graph, and each node is marked as stable or unstable, thus obtaining the labeled mapped directed acyclic graph. This labeling allows for the allocation of appropriate storage and query optimization strategies to each node. For stable nodes, centralized storage and optimized query paths can be implemented, while for unstable nodes, a more distributed storage method can be chosen. Preferably, multiple sample labeled mapped directed acyclic graphs and corresponding multiple sample data partitioning schemes are obtained as scheme identification training data. The scheme identification training data is divided into a training set and a validation set according to a preset ratio. The training set is used to supervise the training of a framework built based on a feedforward neural network. The multiple sample labeled mapped directed acyclic graphs from the validation set are input into the framework to obtain multiple validation data partitioning schemes. The multiple validation data partitioning schemes are compared with the multiple sample data partitioning schemes in the validation set. If the comparison success rate meets the requirements of those skilled in the art, a trained scheme recognizer is obtained. The scheme recognizer is used to identify partitioning schemes in the labeled mapped directed acyclic graph to obtain data partitioning schemes. Finally, based on the data sharding scheme, a mapping routing table is constructed. The table fields include node name, project, stability label, physical database address (IP + port), and storage capacity threshold. For example, a budget of 12 million - project PS-192.168.1.101:3306-500GB is used. The routing table is stored in the gateway module of the dual-architecture data sharding platform for querying.
[0084] Example 2, based on the same inventive concept as the data partitioning method in the foregoing examples, such as... Figure 2 As shown, this application provides a data sharding system. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0085] Spatial reasoning module 11 is used to obtain K project information of the target enterprise, call the long-range spatial reasoning channel in the data sharding dual architecture platform to perform spatial reasoning on the K project information, and obtain the K project information interaction primitive set and the K project information interaction primitive association direction set, where K is a positive integer;
[0086] The mapping directed acyclic graph construction module 12 is configured to construct a mapping directed acyclic graph by taking each item information interaction element in the K item information interaction element set as a node and determining a directed edge according to the K item information interaction element association direction;
[0087] The stability recognition result obtaining module 13 is configured to call a query response processing channel in the data split-database dual-architecture platform, perform state stability time sequence iteration monitoring on each node in the mapping directed acyclic graph, and obtain a plurality of node state stability recognition results;
[0088] The identification module 14 is configured to identify each node of the mapping directed acyclic graph according to the plurality of node state stability recognition results, and obtain an identified mapping directed acyclic graph.
[0089] The mapping routing table construction and storage module 15 is configured to perform split-database scheme recognition based on the identified mapping directed acyclic graph, determine a data split-database scheme, and construct and store a mapping routing table based on the data split-database scheme.
[0090] Further, the spatial reasoning module 11 is configured to perform the following steps:
[0091] The data extraction branch of the long-range spatial reasoning channel is used to perform multi-modal data extraction on the K items of information, and K multi-modal data sets are obtained;
[0092] The data description branch of the long-range spatial reasoning channel is used to generate data universal descriptions for the K multi-modal data sets, and K data universal description text sets are obtained, wherein each data universal description text includes an object, an attribute, a relationship, and a scene.
[0093] The K data universal description text sets are traversed to perform fine-grained spatial reasoning question answering, and K item information interaction element sets and K item information interaction element association direction sets are extracted.
[0094] Further, the spatial reasoning module 11 is configured to perform the following steps:
[0095] K data universal description text sets are used to perform keyword extraction, respectively, and K keyword vector sets are obtained;
[0096] Based on the K keyword vector sets, question dispersion retrieval is performed in a fine-grained spatial reasoning question answering library, and K fine-grained spatial reasoning question group sets are determined.
[0097] The K data universal description text sets are used as reply materials to perform reply analysis on the K fine-grained spatial reasoning question group sets, and K data universal description question reply reliability sets are obtained.
[0098] Filtering the K sets of data general description texts based on the K sets of data general description problem reply reliability, to determine K sets of valid data general description texts;
[0099] Extracting interaction primitives and interaction primitive association directions from the K sets of valid data general description texts by using spaCy, to obtain K sets of project information interaction primitives and K sets of project information interaction primitive association directions.
[0100] Further, the spatial reasoning module 11 is configured to perform the following steps:
[0101] Extracting a first keyword vector from the K sets of keyword vectors, and performing matching in the fine-grained spatial reasoning question and answer library according to the first keyword vector, to obtain a first matched fine-grained spatial reasoning question group;
[0102] Performing question dispersion search on the first matched fine-grained spatial reasoning question group, to determine a first fine-grained spatial reasoning question group;
[0103] Adding the first fine-grained spatial reasoning question group into the K sets of fine-grained spatial reasoning question groups.
[0104] Further, the spatial reasoning module 11 is configured to perform the following steps:
[0105] Randomly extracting M matched fine-grained spatial reasoning questions from the first matched fine-grained spatial reasoning question group, wherein M is a positive integer;
[0106] Enumerating the M matched fine-grained spatial reasoning questions two by two and calculating the dispersion degrees of the enumerated combinations, to obtain a set of enumerated combination dispersion degrees, and performing the same association matched fine-grained spatial reasoning question division and mean calculation on the set of enumerated combination dispersion degrees, to obtain M matched fine-grained spatial reasoning question dispersion mean values;
[0107] Performing question dispersion verification based on the M matched fine-grained spatial reasoning question dispersion mean values, and if the verification fails, obtaining an abnormal matched fine-grained spatial reasoning question and a to-be-supplemented matched fine-grained spatial reasoning question group;
[0108] Taking the abnormal matched fine-grained spatial reasoning question as an index, calculating the dispersion degrees of the matched fine-grained spatial reasoning questions in the first matched fine-grained spatial reasoning question group except for the M matched fine-grained spatial reasoning questions, and performing filtering according to a preset dispersion threshold and ordering according to the dispersion degrees from large to small, to obtain a candidate matched fine-grained spatial reasoning question sequence;
[0109] The candidate matching fine-grained space reasoning question sequence is added into the to-be-supplemented matching fine-grained space reasoning question group in sequence, and question dispersion verification is performed until the verification is passed, and a first fine-grained space reasoning question group is obtained.
[0110] Further, the stability recognition result obtaining module 13 is configured to perform the following steps:
[0111] The state of each node in the mapping directed acyclic graph is monitored in a preset monitoring window according to a preset query response index through a query response processing channel, and a plurality of node state feature sequences are obtained.
[0112] The plurality of node state feature sequences are iteratively interacted with adjacent features within the sequence, and a plurality of interactive node state features are obtained.
[0113] Based on the plurality of interactive node state features, state stability recognition is performed to obtain a plurality of node state stability recognition results.
[0114] Further, the stability recognition result obtaining module 13 is configured to perform the following steps:
[0115] The plurality of first-bit node state features and the plurality of second-bit node state features in the plurality of node state feature sequences are respectively mapped to adjacent sub-features, and a plurality of first adjacent sub-feature deviation sets are obtained.
[0116] The plurality of first adjacent sub-feature deviation sets are normalized and filled into an initially empty adjacent matrix, and a plurality of first adjacent matrices are obtained.
[0117] Based on the plurality of first adjacent matrices, the plurality of second-bit node state features are iteratively interacted, a plurality of first iteratively interacted node state features are obtained, and the plurality of third-bit node state features in the plurality of node state feature sequences are iteratively interacted with adjacent features using the plurality of first iteratively interacted node state features. In this way, the plurality of interactive node state features are obtained.
[0118] Embodiment three, as shown in FIG. 3, is a structural schematic diagram of an example electronic device of the present application. Figure 3 Figure 3 In particular embodiments, bus architecture 300 is represented as a bus 300, which can include any number of interconnecting buses and bridges, and can also include various other circuits and control lines used to interconnect the various elements of the system. Bus 300 can include a bus interface 305 that provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 can be the same component, i.e., a transceiver, providing a unit for communicating with various other systems over a transmission medium.
[0119] The memory 304, as a computer program product, includes computer instructions, which can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the data split library method in the embodiments of the present application. The processor 302 executes various functions and data processing of the computer device by running the software programs, instructions and modules stored in the memory 304, that is, implements the above-mentioned data split library method. The technical features of the above embodiments can be combined in any way. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0120] The above description of disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data library partitioning method, characterized by, The method comprises: Obtaining K project information of a target enterprise, calling a long-range spatial reasoning channel in a data split-database dual architecture platform to perform spatial reasoning on the K project information, and obtaining a K project information interaction element set and a K project information interaction element association direction set, wherein K is a positive integer; Taking each project information interaction element in the K project information interaction element set as a node and determining a directed edge according to the K project information interaction element association direction, a mapping directed acyclic graph is constructed; Calling a query response processing channel in the data split-database dual architecture platform, performing state stability time sequence iteration monitoring on each node in the mapping directed acyclic graph, obtaining a plurality of node state stability identification results, and the node state stability identification result comprises stability and instability. If it is stable, the possibility of inconsistency or performance bottleneck of the current storage path of the corresponding node is lower; According to the plurality of node state stability identification results, each node of the mapping directed acyclic graph is identified, and an identified mapping directed acyclic graph is obtained; Based on the identified mapping directed acyclic graph, a split-database scheme is identified, a data split-database scheme is determined, and a mapping routing table is constructed and stored based on the data split-database scheme.
2. The method of claim 1, wherein, Obtaining K project information of a target enterprise, calling a long-range spatial reasoning channel in a data split-database dual architecture platform to perform spatial reasoning on the K project information, and obtaining a K project information interaction element set and a K project information interaction element association direction set, comprising: Using a data extraction branch of the long-range spatial reasoning channel to perform multi-modal data extraction on the K project information, and obtaining a K multi-modal data set; Using a data description branch of the long-range spatial reasoning channel to generate data universal description for the K multi-modal data set, and obtaining a K data universal description text set, wherein each data universal description text comprises an object, an attribute, a relationship and a scene; Iterating through the K data universal description text set to perform fine-grained spatial reasoning question and answer, and extracting a K project information interaction element set and a K project information interaction element association direction set.
3. The method of claim 2, wherein, Iterating through the K data universal description text set to perform fine-grained spatial reasoning question and answer, and extracting a K project information interaction element set and a K project information interaction element association direction set, comprising: Respectively extracting keywords from the K data universal description text set to obtain a K keyword vector set; Based on the K keyword vector set, performing problem dispersion retrieval in a fine-grained spatial reasoning question and answer library to determine a K fine-grained spatial reasoning question group set; Taking the K data universal description text set as a reply material, replying and analyzing the K fine-grained spatial reasoning question group set to obtain a K data universal description question and reply reliability set; Based on the K data universal description question and reply reliability set, screening the K data universal description text set to determine a K effective data universal description text set; Interactives and interactive association directions of the K sets of valid data general description texts are extracted by using spaCy to obtain K sets of project information interactive elements and K sets of project information interactive association directions.
4. The method of claim 3, wherein, Based on the K sets of keyword vectors, question dispersion retrieval is performed in the fine-grained space reasoning question and answer library to determine K sets of fine-grained space reasoning question groups, including: A first keyword vector is extracted from the K sets of keyword vectors, and a first matched fine-grained space reasoning question group is obtained by matching the first keyword vector in the fine-grained space reasoning question and answer library. Question dispersion retrieval is performed on the first matched fine-grained space reasoning question group to determine a first fine-grained space reasoning question group. The first fine-grained space reasoning question group is added to the K sets of fine-grained space reasoning question groups.
5. The method of claim 4, wherein, Question dispersion retrieval is performed on the first matched fine-grained space reasoning question group to determine a first fine-grained space reasoning question group, including: M matched fine-grained space reasoning questions are randomly extracted from the first matched fine-grained space reasoning question group, where M is a positive integer. The M matched fine-grained space reasoning questions are enumerated two by two, and dispersion calculation is performed on the enumerated combinations to obtain a set of enumerated combination dispersion values. The set of enumerated combination dispersion values is divided into the same associated matched fine-grained space reasoning questions and the mean value is calculated to obtain M matched fine-grained space reasoning question dispersion mean values. Based on the M matched fine-grained space reasoning question dispersion mean values, question dispersion verification is performed. If the verification fails, an abnormal matched fine-grained space reasoning question and a set of to-be-supplemented matched fine-grained space reasoning questions are obtained. The candidate matched fine-grained space reasoning question sequence is obtained by sequentially adding the candidate matched fine-grained space reasoning question sequence to the to-be-supplemented matched fine-grained space reasoning question group, and performing question dispersion verification until the verification passes to obtain the first fine-grained space reasoning question group. The data library dual-architecture platform is called to query the response processing channel, and state stability time sequence iteration monitoring is performed on each node in the mapping directed acyclic graph to obtain a plurality of node state stability identification results, including:
6. The method of claim 1, wherein, Through the query response processing channel, state monitoring is performed on each node in the mapping directed acyclic graph within a preset monitoring window according to a preset query response index to obtain a plurality of node state feature sequences. The plurality of node state feature sequences are sequentially iteratively interacted to obtain a plurality of interactive node state features. Based on the plurality of interactive node state features, state stability identification is performed to obtain a plurality of node state stability identification results. The plurality of node state feature sequences are sequentially iteratively interacted to obtain a plurality of interactive node state features, including:
7. The method of claim 6, wherein the data is stored in the first database and the second database in a manner such that the data is accessible from either the first database or the second database. Adjoint feature mapping deviation identification is respectively performed on the first and second bit node state features in the plurality of node state feature sequences, to obtain a plurality of first adjoint feature deviation degree sets; The plurality of first adjoint feature deviation degree sets are normalized and filled into an initially empty adjoint matrix, to obtain a plurality of first adjoint matrices; Based on the plurality of first adjoint matrices, second bit node state features are respectively iteratively interacted to obtain a plurality of first iteratively interacted node state features, and third bit node state features in the plurality of node state feature sequences are iteratively interacted based on the plurality of first iteratively interacted node state features, and so on, to obtain the plurality of interacted node state features.
8. A data partitioning system, comprising: The system is used to perform a data library splitting method according to any one of claims 1-7, and the system comprises: A spatial reasoning module is configured to obtain K pieces of project information of a target enterprise, and call a long-range spatial reasoning channel in a data library splitting double-architecture platform to perform spatial reasoning on the K pieces of project information, to obtain a K-piece project information interaction primitive set and a K-piece project information interaction primitive association direction set, where K is a positive integer; A mapping directed acyclic graph construction module is configured to take each project information interaction primitive in the K-piece project information interaction primitive set as a node, determine a directed edge according to the K-piece project information interaction primitive association direction, and construct a mapping directed acyclic graph; A stability identification result obtaining module is configured to call a query response processing channel in the data library splitting double-architecture platform to perform state stability time sequence iteration monitoring on each node in the mapping directed acyclic graph, to obtain a plurality of node state stability identification results, where the node state stability identification result includes stability and instability, and if the node state stability identification result is stable, the possibility of inconsistency or performance bottleneck of a current storage path of the corresponding node is lower; An identification module is configured to identify each node in the mapping directed acyclic graph according to the plurality of node state stability identification results, to obtain an identified mapping directed acyclic graph; A mapping routing table construction and storage module is configured to perform library splitting scheme identification based on the identified mapping directed acyclic graph, determine a data library splitting scheme, and construct and store a mapping routing table based on the data library splitting scheme.
9. An electronic device, comprising: The electronic device comprises: A memory configured to store executable instructions; A processor configured to execute the executable instructions stored in the memory, to implement a data library splitting method according to any one of claims 1-7.
10. A computer program product comprising computer instructions, characterized in that, The computer instructions are executed by the processor to implement a data library splitting method according to any one of claims 1-7.
Citation Information
Patent Citations
Database dividing method and device, electronic equipment and storage medium
CN117709903A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and automation
US20250258852A1