Optimization control method and device for graph cleaning rule matching, equipment and medium

By initializing the memory pool and bucket partitioning memory slots during the matching process of graph cleaning rules, the memory pressure and frequent memory operations are solved, and efficient memory reuse and performance improvement are achieved.

CN120216203APending Publication Date: 2025-06-27SHENZHEN INST OF COMPUTING SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510596652.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing technology has huge memory pressure in the matching process of graph cleaning rules, resulting in frequent memory application and release operations, increasing memory management overhead and affecting overall performance.

Method used

By initializing continuous memory areas as memory pools, when the graph cleaning rule matching task is submitted, a memory segment is directly selected from the memory pool, and the memory segments are evenly divided according to the number of buckets in the instance bucket splitter to obtain the memory slots with the corresponding bucket number. For any bucket, after the bucket obtains the instance data, the data is written to the corresponding memory slot. After the memory slot is full, the data is added to the disk file and the write pointer is updated for subsequent direct writing.

Benefits of technology

By introducing memory buffer pool and memory slot mechanisms, efficient memory reuse is achieved, frequent memory application and release operations are avoided, and memory management overhead is significantly reduced, thereby improving the overall performance of graph cleaning rules matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216203A_ABST
    Figure CN120216203A_ABST
Patent Text Reader

Abstract

The invention relates to an optimization control method and device for graph cleaning rule matching, equipment and a medium. According to the method, a continuous memory area is initialized to serve as a memory pool, a memory segment is directly selected from the memory pool when a graph cleaning rule matching task is submitted, the memory segment is evenly divided according to the bucket dividing number of bucket dividing performed by an instance bucket dividing device to obtain memory grooves with the corresponding number, and after instance data is obtained through bucket dividing, the graph cleaning rule matching task is submitted. And writing the instance data into the memory slots of the corresponding buckets, adding the data in the memory slots into the disk file after the memory slots are full, updating the write-in pointers of the memory slots after the addition is completed, and completely writing the remaining data in all the memory slots into the disk file after all the data are bucket-divided. By introducing a memory buffer pool and a memory slot mechanism, efficient multiplexing of the memory is realized, frequent memory application and release operations are avoided, and the memory management overhead is remarkably reduced, so that the overall performance of graph cleaning rule matching is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technology, and in particular, to an optimized control method, device, equipment and medium for graph cleaning rule matching. Background Art

[0002] The graph cleaning rule is an important graph data cleaning tool, and its mathematical form can be expressed as: φ = Q[x0, y0](X → p0), where Q is a double-star graph pattern used to identify relevant entities in the graph; X is a conjunction of a set of predicates defined on leaf nodes or central nodes; p0 is a predicate defined on leaf nodes or central nodes; X → p0 can represent the association and dependency relationships between entities in the graph.

[0003] The graph data cleaning workflow based on graph cleaning rules is as follows: (1) In the rule discovery stage, graph cleaning rules are enumerated according to user configuration, and the enumerated rules are matched in the data graph sample input by the user. Rules with both support and confidence higher than the user-set threshold will be added to the rule set and output; (2) In the error detection stage, the graph G and the rule set Σ are accepted as inputs, and these rules are matched in the data graph. Entities that do not satisfy the rule dependencies will be output as conflicts; (3) In the error repair stage, the graph G, the rule set Σ, and the conflict data are accepted as inputs, the errors in the graph data are repaired, and the error detection is performed again in the second stage. When there is no conflict data output in the second stage, the process terminates.

[0004] In the above workflow, the matching of graph cleaning rules runs through the graph data cleaning process. Therefore, the optimization of graph cleaning rule matching is crucial for the overall performance optimization of the graph data cleaning system.

[0005] The existing optimization techniques for graph cleaning rule matching can be classified into the following two categories: (1) Graph cleaning rule combined matching technology. This type of technology first disassembles each graph cleaning rule into two single-star components, then combines the single-star components with shared substructures into a single-star bundle, and finally performs matching with the single-star bundle as a whole; (2) Matching bucketing technology. This type of technology aims to avoid the pairwise enumeration with a complexity of O(N 2 ) (N is the number of matches of single-star components) in the double-star matching enumeration stage. Its main idea is to divide the matching single-star component instances into K buckets according to the predicates of the graph cleaning rules. In this way, only when two single-star component instances are in the same bucket can they be combined into a graph cleaning rule instance. The matching bucketing strategy reduces the complexity of double-star matching enumeration from O(N 2 ) to O( ), where , which greatly reduces the computational overhead required for enumeration.

[0006] To cope with the huge memory pressure brought by graph cleaning rule matching, the existing system has to frequently apply for and release memory. Especially for the matching bucketing technology, to achieve fine-grained memory management at the bucket level, the system is further required to allocate and release memory in units of buckets, which undoubtedly increases the system overhead in memory management.

[0007] Therefore, how to improve memory operations during graph cleaning rule matching to enhance memory operation efficiency has become an urgent problem to be solved. Summary of the Invention

[0008] In view of this, the embodiments of the present application provide an optimized control method, device, equipment and medium for graph cleaning rule matching to solve the problem of how to improve memory operations during graph cleaning rule matching to enhance memory operation efficiency.

[0009] In a first aspect, the embodiments of the present application provide an optimized control method for graph cleaning rule matching, including: Initialize a continuous memory area as a memory pool, and directly select a memory segment from the memory pool when a graph cleaning rule matching task is submitted; According to the number of buckets obtained by bucketing the matching instances retrieved from the preset graph data by an instance bucketer, evenly divide the memory segment to obtain memory slots corresponding to the number of buckets; For any bucket, after the bucket obtains instance data, write the instance data into the memory slot corresponding to the bucket; For any memory slot, after the memory slot is full, append the data in the memory slot to the corresponding file on the disk, and update the write pointer of the memory slot after the append is completed. The write pointer is used to indicate that subsequent data is directly written; After all data is bucketed, write all the remaining data in all memory slots into the corresponding file on the disk.

[0010] In a second aspect, an optimized control device for graph cleaning rule matching according to the embodiments of the present application includes: An initialization module, configured to initialize a continuous memory area as a memory pool, and directly select a memory segment from the memory pool when a graph cleaning rule matching task is submitted; A memory slot construction module, configured to evenly divide the memory segment according to the number of buckets obtained by bucketing the matching instances retrieved from the preset graph data by an instance bucketer to obtain memory slots corresponding to the number of buckets; A data writing module, configured to, for any bucket, write the instance data into the memory slot corresponding to the bucket after the bucket obtains the instance data; A repeated persistence module, which is used for any memory slot, after the memory slot is full, append the data in the memory slot to the corresponding file on the disk, and update the write pointer of the memory slot after the append is completed, where the write pointer is used to indicate that subsequent data is directly written; A full persistence module, which is used to write all the remaining data in all memory slots to the corresponding file on the disk after all data is bucketed.

[0011] In a third aspect, an embodiment of the present application provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the optimized control method for graph cleaning rule matching as described in the first aspect.

[0012] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the optimized control method for graph cleaning rule matching as described in the first aspect.

[0013] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: In the present application, a continuous memory area is initialized as a memory pool. When a graph cleaning rule matching task is submitted, a memory segment is directly selected from the memory pool, and the memory segment is evenly divided according to the number of buckets obtained by bucketizing the matching instances obtained from the preset graph data by an instance bucketizer, to obtain memory slots corresponding to the number of buckets. For any bucket, after the bucket obtains instance data, the instance data is written into the memory slot corresponding to the bucket. For any memory slot, after the memory slot is full, the data in the memory slot is appended to the corresponding file on the disk, and the write pointer of the memory slot is updated after the append is completed, where the write pointer is used to indicate that subsequent data is directly written. After all data is bucketed, all the remaining data in all memory slots is written to the corresponding file on the disk. By introducing a memory buffer pool and a memory slot mechanism, efficient reuse of memory is achieved, frequent memory application and release operations are avoided, the memory management overhead is significantly reduced, and thus the overall performance of graph cleaning rule matching is effectively improved. Brief Description of the Drawings

[0014] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 It is a schematic diagram of an application environment of an optimization control method for graph cleaning rule matching provided in the first embodiment of the present application; Figure 2 It is a schematic flowchart of an optimization control method for graph cleaning rule matching provided in the second embodiment of the present application; Figure 3 It is a data example diagram of graph cleaning rules provided in the second embodiment of the present application; Figure 4 It is a schematic flowchart of a graph data cleaning workflow provided in the second embodiment of the present application; Figure 5 It is a data example diagram of a graph cleaning rule combination provided in the second embodiment of the present application; Figure 6 It is a schematic flowchart of memory management and concurrency control provided in the second embodiment of the present application; Figure 7 It is a schematic flowchart of an optimization control method for graph cleaning rule matching provided in the third embodiment of the present application; Figure 8 It is a schematic flowchart of an optimization control method for graph cleaning rule matching provided in the fourth embodiment of the present application; Figure 9 It is a schematic diagram of the formation of a conditional compact table provided in the fourth embodiment of the present application; Figure 10 It is a schematic diagram of the matching reconstruction of a conditional compact table provided in the fourth embodiment of the present application; Figure 11 It is a schematic structural diagram of an optimization control device for graph cleaning rule matching provided in the fifth embodiment of the present application; Figure 12 It is a schematic structural diagram of a computer device provided in the sixth embodiment of the present application. Detailed Description of the Invention

[0016] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0017] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0018] It should also be understood that the term "and / or" as used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0019] As used in the specification of this application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0020] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for differentiating descriptions and cannot be understood as indicating or implying relative importance.

[0021] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "include but not be limited to", unless otherwise specifically emphasized in other ways.

[0022] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0023] In order to illustrate the technical solution of this application, specific embodiments will be used for illustration below.

[0024] An optimization control method for graph cleaning rule matching provided in the first embodiment of this application can be applied in, for example Figure 1In the application environment, the client communicates with the server. The optimization control method is executed on the server, and the client sends a corresponding start instruction to the server to start the graph cleaning rule matching. Graph data can be stored in the server, and the graph data can be directly called during graph cleaning rule matching. The server is configured with a storage disk, which is used to provide memory support for graph cleaning rule matching and persistent storage of files such as instance data.

[0025] Among them, the client includes but is not limited to computer devices such as a palm computer, a desktop computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud terminal device, and a personal digital assistant (PDA). The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0026] See Figure 2 , which is a schematic flowchart of an optimization control method for graph cleaning rule matching provided in the second embodiment of the present application. The above optimization control method for graph cleaning rule matching can be applied to Figure 1 the server in, and the server performs optimization control during graph cleaning rule matching, specifically for the optimization control of data storage during the graph cleaning rule matching process.

[0027] Such as Figure 2 shown, the optimization control method for graph cleaning rule matching may include the following steps: Step S201, initialize a continuous memory area as a memory pool, and directly select a memory segment from the memory pool when the graph cleaning rule matching task is submitted.

[0028] In this embodiment, the memory storage area is characterized as the area contained in the memory of the server. Select a continuous memory area from the memory and initialize this memory area, that is, clear this memory area as the memory pool, so as to facilitate the storage of intermediate data during all subsequent graph cleaning rule matching processes.

[0029] The process of graph cleaning rule matching in the present application is as follows: Figure 3 shows an example of a graph cleaning rule. Rule will first match the double-star graph pattern Q in graph G, and then assert that in each instance of the double-star graph pattern, if the machine learning model Ms It is considered that if the conferences of two paper nodes are the same as the author names, and the category of one of the papers is CS, then the category of the other paper should also be CS. This mathematical form based on logical reasoning ensures the correctness and interpretability of the graph data cleaning results. In addition, the graph cleaning rules are also compatible with machine learning models in the form of predicates, so the data cleaning results also have strong generalization ability. Due to its strong theoretical guarantee, more and more graph data cleaning systems choose graph cleaning rules as the theoretical basis.

[0030] Figure 4 is a schematic diagram of the graph data cleaning workflow. (1) In the rule discovery stage, graph cleaning rules are enumerated according to user configuration, and the enumerated rules are matched in the data graph sample input by the user. Rules with both support and confidence higher than the user-set threshold will be added to the rule set and output. (2) The error detection stage takes the graph G and the rule set Σ as inputs and matches these rules in the data graph. Entities that do not satisfy the rule dependencies will be output as conflicts. (3) The error repair stage takes the graph G, the rule set Σ, and the conflict data as inputs, repairs the errors in the graph data, and returns to the second stage to perform error detection again. When no conflict data is output in the second stage, the process terminates.

[0031] In the above workflow, the matching of graph cleaning rules runs through the graph data cleaning process. Therefore, the optimization of the matching of graph cleaning rules is crucial for the overall performance optimization of the graph data cleaning system.

[0032] The matching of this graph cleaning rule usually includes two stages: (1) Single-star graph pattern matching. In this stage, the double-star pattern of the graph cleaning rule is disassembled into two single-star components for separate matching to avoid excessive intermediate results caused by directly materializing the double-star pattern. (2) Double-star matching enumeration. In this stage, pairs of the two single-star patterns corresponding to the graph cleaning rule are enumerated and checked to see if the matches satisfy the predicate dependencies defined by the rule. Instance pairs that do not satisfy the dependencies will be regarded as data conflicts.

[0033] Each graph cleaning rule is disassembled into two single-star components, and then the single-star components with shared substructures are combined into single-star bundles, and finally the single-star bundles are matched as a whole. Figure 5 shows a data example of graph cleaning rule combination. In this example, the double-star pattern of the rule can be disassembled into and two single-star components. These two single-star components share three nodes x0, x1, x2 and two edges x0→x1, x0→x2. By merging these two single-star components and performing unified matching, the system can eliminate the repeated matching of this common substructure.

[0034] When the graph cleaning rule matching task is submitted, a memory segment is directly selected from the memory pool. Of course, this memory segment is an unoccupied memory segment, so that the corresponding example data can be stored. For example, a memory segment with a starting address of p and a size of s is pre-allocated from the memory pool. Since this memory is directly taken from the memory pool, the overhead of the traditional malloc() function is avoided, thus eliminating the time-consuming of memory application.

[0035] Step S202: According to the number of buckets obtained by the instance bucketizer for bucketing the matching instances obtained from the preset graph data, the memory segment is evenly divided to obtain memory slots corresponding to the number of buckets.

[0036] Among them, when performing graph cleaning rule matching, bucketing is performed on the preset graph data, and the number of corresponding buckets determines the number of memory slots, that is, one bucket corresponds to one memory slot.

[0037] Figure 6 This is a schematic flowchart of memory management and concurrency control provided by the second embodiment of the present application. For example, when pattern matching is performed on a graph through a pattern matcher, the obtained matching instances enter the instance splitter for bucketing. If there are N buckets, in order to support the bucketing operation, this memory segment will be further divided into N memory slots of uniform size, that is, the i-th bucket corresponds to a memory slot with a starting address of (i - 1)s / N and a size of s / N.

[0038] Step S203: For any bucket, after the instance data is obtained in the bucket, the instance data is written into the memory slot corresponding to the bucket.

[0039] Among them, after the instance data is obtained in the corresponding bucket, the instance data is written into the memory slot corresponding to the bucket. Of course, before that, the memory slot can be initialized, and after the initialization is completed, the example data is written into the memory slot.

[0040] Step S204: For any memory slot, after the memory slot is full, the data in the memory slot is appended to the corresponding file on the disk, and the write pointer of the memory slot is updated after the append is completed.

[0041] Among them, when a certain memory slot is full, the system does not immediately call the free() function to release the memory, but appends the data in the slot to the corresponding file on the disk and only updates the write pointer of the memory slot, and subsequent data can be directly written.

[0042] The write pointer is used to indicate that subsequent data is directly written. Since the write pointer of this memory slot is continuously updated as the stored data increases or decreases during the storage process, data can be stored continuously. After the data in the memory slot is flushed to the corresponding file on the disk for persistence, the content in this memory slot can be cleared. Therefore, the write pointer of this memory slot can be set to the initial position to guide data writing. This process is looped until all data is bucketed.

[0043] Step S205, after all data is bucketed, all the remaining data in all memory slots is written into the corresponding file on the disk.

[0044] Among them, after all data is bucketed, the system will write all the remaining data in all memory slots to the disk at one time to ensure that all data is persisted. Therefore, by introducing the memory buffer pool and memory slot mechanism, efficient reuse of memory is achieved, frequent memory application and release operations are avoided, the memory management overhead is significantly reduced, and thus the overall performance of graph cleaning rule matching is effectively improved.

[0045] In the embodiment of the present application, a continuous memory area is initialized as a memory pool. When a graph cleaning rule matching task is submitted, a memory segment is directly selected from the memory pool. According to the number of buckets obtained by bucketing the matching instances retrieved from the preset graph data by an instance bucketizer, the memory segment is evenly divided to obtain memory slots corresponding to the number of buckets. For any bucket, after the bucket obtains instance data, the instance data is written into the memory slot corresponding to the bucket. For any memory slot, after the memory slot is full, the data in the memory slot is appended to the corresponding file on the disk. After the append is completed, the write pointer of the memory slot is updated. The write pointer is used to indicate that subsequent data is directly written. After all data is bucketed, all the remaining data in all memory slots is written into the corresponding file on the disk. By introducing the memory buffer pool and memory slot mechanism, efficient reuse of memory is achieved, frequent memory application and release operations are avoided, the memory management overhead is significantly reduced, and thus the overall performance of graph cleaning rule matching is effectively improved.

[0046] See Figure 7 , which is a schematic flowchart of an optimization control method for graph cleaning rule matching provided in Embodiment 3 of the present application. As Figure 7 shown, in the above step S202, according to the number of buckets obtained by bucketing the matching instances retrieved from the preset graph data by an instance bucketizer, evenly dividing the memory segment to obtain memory slots corresponding to the number of buckets may include the following steps: Step S701, obtain matching instances, calculate the bucket identifier to which the matching instances belong, and trigger the instance data corresponding to the matching instances to be written into the bucket corresponding to the bucket identifier.

[0047] Step S702: Increase the memory occupancy count of the bucket and detect whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instance.

[0048] Step S703: If it is detected that the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instance, determine the write offset according to the memory occupancy count, and write the instance data corresponding to the matching instance into the memory slot corresponding to the bucket according to the write offset.

[0049] In this embodiment, the corresponding bucket is determined by calculating the bucket identifier, and by calculating the similarity of the matching instances, the bucket to which the similar instances belong is determined, thereby determining the bucket identifier.

[0050] The memory occupancy count is used to record the memory occupancy of the corresponding bucket, and can be used to determine the write offset when writing data next time, so as to know the data writing, and can also be used to determine whether the current memory slot is full, thereby determining the actual write offset.

[0051] Optionally, after detecting whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instance, it further includes: If it is detected that the memory slot corresponding to the bucket is not sufficient to accommodate the instance data corresponding to the matching instance, determine that the memory slot is full, and execute the step of appending the data in the memory slot to the corresponding file on the disk; After updating the write pointer of the memory slot after the append is completed, it further includes: Write the instance data corresponding to the matching instance into the memory slot corresponding to the bucket according to the write pointer.

[0052] Wherein, if the memory slot is full, it is necessary to execute the step of appending the data in the memory slot to the corresponding file on the disk, so as to write the instance data corresponding to the newly arrived matching instance into the memory slot through the updated write pointer.

[0053] Optionally, increasing the memory occupancy count of the bucket includes: Obtain the current occupancy count of the bucket and the instance size of the instance data corresponding to the matching instance; Based on the atomic variable, use the instance size to increase the current occupancy count to obtain the updated occupancy count, and according to the updated occupancy count, execute the detection of whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instance, and return the current occupancy count; Determining the write offset according to the memory occupancy count includes: Determine the write offset with the returned current occupancy count.

[0054] Wherein, as Figure 6As shown, the update of the occupancy technique is achieved through the atomic variable method, which essentially accumulates the instance size. This atomic variable method can improve the calculation efficiency of the write offset, thereby enhancing the storage efficiency.

[0055] Optionally, after appending the data in the memory slot to the corresponding file on the disk, it further includes: Reset the current occupancy count formed by the atomic variable to 0 and restart the write thread blocked by the condition variable; After updating the write pointer of the memory slot after the append is completed, it further includes: Based on the write thread, restart the bucket writing operation according to the write pointer.

[0056] Among them, the specific implementation steps are as follows: (1) The system first completes the initialization of the memory pool and memory slots to prepare for subsequent data storage; (2) The instance bucketizer first calculates the bucket ID to which the matching instance belongs and triggers the instance bucket entry operation; (3) When an instance attempts to enter the bucket, it atomically increases the memory occupancy count of the corresponding bucket (implemented through an atomic variable, essentially accumulating the instance size). This atomic operation returns the count value before the update, which is the write offset of the current instance in the memory slot; (4) Subsequently, the condition variable associated with this bucket checks whether there is enough remaining space in the current memory slot to accommodate the instance. If the space is sufficient (i.e., the memory slot occupancy is less than or equal to the memory slot size s / N), the instance is written to the memory slot according to the calculated offset; otherwise, if the memory slot is full (as shown in Figure 6 memory slot 1 (full memory) in the figure), the system immediately triggers an I / O operation to persist the data in this slot to the corresponding file on the disk. At this time, the atomic variable of this bucket is reset to 0, and all threads blocked by the condition variable are awakened to re-attempt the bucket entry operation (return to (3)); (5) When all data has completed entering the bucket, the system will uniformly write the remaining data in all memory slots to the disk to ensure data integrity. As described above, this concurrent control algorithm uses atomic operations to reduce the use of mutex locks, significantly reducing the performance overhead caused by lock contention, thereby effectively enhancing the concurrent processing ability of the system.

[0057] During the parallel rule matching process, multiple threads perform memory read and write activities simultaneously. To ensure the content consistency of the memory critical section in a concurrent environment, mutex locks are usually relied on to implement concurrent control. However, the extensive use of mutex locks significantly reduces the performance of the parallel system. In the embodiments of the present application, an efficient concurrent bucketing mechanism is implemented by combining atomic variables and condition variables. The atomic variables are used to track the memory occupancy of each bucket and atomically obtain the write offset, avoiding lock contention. The condition variables are used to monitor the remaining space in the memory slot. When the slot is full, it triggers data flushing to disk and wakes up the blocked threads to retry bucketing. This design reduces the use of traditional mutex locks, significantly reduces the overhead of concurrent control, and improves the concurrent processing performance of the system.

[0058] See Figure 8 , which is a schematic flowchart of an optimization control method for graph cleaning rule matching provided in the fifth embodiment of the present application. As Figure 8 shown, if all matching instances are single-star beam structure tables, the single-star beam structure table includes a central node and at least one leaf node. The central node and each leaf node respectively correspond to an attribute, and each attribute corresponds to a data value under any instance data. After writing the instance data into the memory slot of the corresponding bucket in step S203, the following steps may be included: Step S801, compress the matching instances in the memory slot of the bucket, and use the union of the attributes that make up all single-star beam structure tables in all matching instances as the header of the conditional compact table.

[0059] Among them, by compressing the data, a compact data structure can be formed, which can reduce the memory pressure and frequent disk I / O to a certain extent.

[0060] At this time, for the instance data in the memory slot corresponding to the bucket, it needs to be stored after compression. The instance data is structured data and is stored in the form of a table. As Figure 9 shown, it is a schematic diagram of the formation of a conditional compact table provided in the fourth embodiment of the present application. The two tables in the dotted line are two single-star beam structure tables respectively, and the conditional compact table extracts the union of all single-star component attributes that make up the single-star beam as the header.

[0061] Step S802, store all matching instances with the same central node as the root in the same row of the conditional compact table to obtain the compressed conditional compact table.

[0062] Among them, since all leaf nodes of the single-star beam share the same central node, the conditional compact table stores all candidate matching results with the same central node as the root in the same row, and each column corresponds to the matching result of a leaf node.

[0063] The above structure avoids the redundancy of storing each candidate match separately in the traditional table structure, thus greatly saving storage space.

[0064] As Figure 9 shown, assume that table and table are the matching results of two single-star components. The conditional compact table is used to store the matches of the single-star bundles formed by their combination. The header attributes of the conditional compact table are the union of the header attributes of table and table , so the storage of duplicate column data can be effectively avoided (i.e., , , ). At the same time, for the matching results with the same central node, the four matching records in table are compressed into two rows in the conditional compact table, thus also reducing the storage of duplicate row data.

[0065] Optionally, after obtaining the compressed conditional compact table, it further includes: Obtaining the attributes of each matching instance and the predicate dependency conditions that satisfy the corresponding attributes, and forming conditional annotations for the corresponding matching instances according to the attributes and predicate dependency conditions; Corresponding the conditional annotations of all matching instances with the compressed conditional compact table, so that the corresponding matching instances can be reconstructed according to the conditional annotations and the compressed conditional compact table.

[0066] Among them, the system maintains a set of conditional annotations for each single-star component. These annotations define the attributes that the single-star component needs to materialize and the predicate dependency conditions that the attributes must satisfy. When it is necessary to reconstruct the matches of a specific single-star component, the system only needs to filter according to the corresponding conditional annotations to efficiently extract the candidate matches that meet the conditions. Figure 10 shows an example of match reconstruction. In this example, the system maintains conditional annotations for the single-star component , including the attributes that need to be materialized and the attribute-related predicates . Based on these conditional annotations, the system can quickly filter out the corresponding two rows (not satisfying the predicate constraints), and , the corresponding two columns (non-materialized attributes), so as to efficiently reconstruct the candidate matching results of the target single-star component.

[0067] During the graph cleaning rule matching process, a large amount of intermediate results will be generated, and the size of these intermediate results may far exceed the memory capacity of a single machine. Therefore, the matching process will frequently trigger disk reads and writes, bringing a huge I / O overhead. In the embodiments of the present application, the matching results of the same central node are compressed into one line, and through the attribute intersection header and conditional annotation mechanism, the storage redundancy is further reduced. During data reconstruction, conditional annotations can guide the system to quickly filter and extract candidate matches of specific single-star components, avoiding full-table scans. Through structured compression and conditional reconstruction, the present invention effectively alleviates the memory and I / O bottlenecks faced by traditional data structures when storing single-star bundle matches.

[0068] For the present application, for the dataset Internet Movie Database (IMDB), the number of its nodes is 5.2 million, and the number of edges is 48.2 million. The solution of the present application has achieved a 13.06-fold system acceleration and a 5.71-fold data compression compared with the latest graph cleaning rule matching system.

[0069] Corresponding to the optimization control method for graph cleaning rule matching in the above embodiments, Figure 11 The structural block diagram of the optimization control device for graph cleaning rule matching provided in the sixth embodiment of the present application is shown. The above optimization control device for graph cleaning rule matching can be applied to Figure 1 the server in. When the server performs graph cleaning rule matching, it executes optimization control, specifically for the optimization control of data storage during the graph cleaning rule matching process. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0070] See Figure 11 , the optimization control device for graph cleaning rule matching includes: An initialization module 1101, configured to initialize a continuous memory area as a memory pool, and directly select a memory segment from the memory pool when a graph cleaning rule matching task is submitted; A memory slot construction module 1102, configured to evenly divide the memory segment according to the number of buckets obtained by binning the matching instances obtained from the preset graph data by an instance binning device, to obtain memory slots corresponding to the number of buckets; A data writing module 1103, configured to, for any bucket, after the bucket obtains instance data, write the instance data into the memory slot corresponding to the bucket; A repeated persistence module 1104, configured to, for any memory slot, after the memory slot is full, append the data in the memory slot to a corresponding file on the disk, and update the write pointer of the memory slot after the append is completed. The write pointer is used to indicate that subsequent data is directly written; The all-persistence module 1105 is used to write all the remaining data in all memory slots to the corresponding files on the disk after all data has been bucketed.

[0071] Optionally, the optimization control device for graph cleaning rule matching further includes: The bucket identification module is used to evenly divide the memory segment according to the number of buckets obtained by bucketing the matching instances retrieved from the preset graph data by the instance bucketizer, obtain the matching instances after obtaining the memory slots corresponding to the number of buckets, calculate the bucket identification to which the matching instances belong, and trigger the writing of the instance data corresponding to the matching instances to the bucket corresponding to the bucket identification; The bucket detection module is used to increase the memory occupancy count of the bucket and detect whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instances; The first bucket writing module is used to, if it is detected that the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instances, determine the writing offset according to the memory occupancy count, and write the instance data corresponding to the matching instances to the memory slot corresponding to the bucket according to the writing offset.

[0072] Optionally, the optimization control device for graph cleaning rule matching further includes: The disk writing module is used to, after detecting whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instances, if it is detected that the memory slot corresponding to the bucket is not sufficient to accommodate the instance data corresponding to the matching instances, determine that the memory slot is full, and execute appending the data in the memory slot to the corresponding file on the disk; The second bucket writing module is used to update the writing pointer of the memory slot after the appending is completed, and write the instance data corresponding to the matching instances to the memory slot corresponding to the bucket according to the writing pointer.

[0073] Optionally, the bucket detection module includes: The instance size acquisition unit is used to acquire the current occupancy count of the bucket and the instance size of the instance data corresponding to the matching instances; The bucket detection unit is used to, based on the atomic variable, use the instance size to increase the current occupancy count to obtain the updated occupancy count, and execute detecting whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instances according to the updated occupancy count, and return the current occupancy count; The first bucket writing module includes: The first writing offset determination unit is used to determine the writing offset with the returned current occupancy count.

[0074] Optionally, the optimization control device for graph cleaning rule matching further includes: A thread reset module, which is used to reset the current occupancy count formed by the atomic variable to 0 after appending the data in the memory slot to the corresponding file on the disk, and restart the write thread blocked by the condition variable; A write restart module, which is used to update the write pointer of the memory slot after the append is completed, and then, based on the write thread and according to the write pointer, restart the bucket writing operation.

[0075] Optionally, if all matching instances are single-star beam structure tables, where a single-star beam structure table includes a central node and at least one leaf node, the central node and each leaf node respectively correspond to an attribute, and each attribute corresponds to a data value under any instance data, the optimization control device for graph cleaning rule matching further includes: A compressed table header module, which is used to compress the data of the matching instances in the memory slot of the bucket after writing the instance data into the corresponding bucket, and use the union of the attributes that make up all single-star beam structure tables in all matching instances as the header of the conditional compact table; A compressed table row module, which is used to store all matching instances with the same central node as the root in the same row of the conditional compact table to obtain the compressed conditional compact table.

[0076] Optionally, the optimization control device for graph cleaning rule matching further includes: A conditional annotation module, which is used to obtain the attributes of each matching instance and the predicate dependency conditions that satisfy the corresponding attributes after obtaining the compressed conditional compact table, and form the conditional annotation of the corresponding matching instance according to the attributes and the predicate dependency conditions; An instance reconstruction module, which is used to correspond the conditional annotations of all matching instances to the compressed conditional compact table, so that the corresponding matching instances can be reconstructed according to the conditional annotations and the compressed conditional compact table.

[0077] It should be noted that, for the information interaction, execution process, etc. among the above modules, units, and subunits, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought thereby can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0078] Figure 12 This is a schematic structural diagram of a computer device provided in Embodiment 7 of the present application. As Figure 12 shown, the computer device in this embodiment includes: at least one processor ( Figure 12 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, it implements the steps in any of the above optimization control methods for graph cleaning rule matching or the optimization control method embodiment for graph cleaning rule matching.

[0079] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 12 merely examples of computer devices, which do not constitute a limitation on computer devices. A computer device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include a network interface, a display screen, an input device, etc.

[0080] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0081] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit and the external storage device of the computer device. The memory is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of a computer program. The memory may also be used to temporarily store data that has been output or will be output.

[0082] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0083] All or part of the processes in the above-mentioned method embodiments of this application can also be completed by a computer program product. When the computer program product runs on a computer device, the computer device can be made to execute the steps in the above-mentioned method embodiments.

[0084] In the above-mentioned embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0085] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0086] In the embodiments provided in this application, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0087] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0088] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. An optimization control method for graph cleaning rule matching, characterized in that: include: Initialize a continuous memory area as a memory pool, and when a graph cleaning rule matching task is submitted, directly select a memory segment from the memory pool; According to the number of buckets into which the instance bucketizer buckets the matching instances obtained from the preset graph data, the memory segment is evenly divided to obtain memory slots corresponding to the number of buckets; For any bucket, after the instance data is obtained in the bucket, the instance data is written into the memory slot corresponding to the bucket; For any memory slot, after the memory slot is full, the data in the memory slot is appended to the file corresponding to the disk, and after the appending is completed, the write pointer of the memory slot is updated, and the write pointer is used to indicate that subsequent data is directly written; After all data is bucketed, all remaining data in all memory slots are written to the corresponding files on the disk.

2. The optimization control method for graph cleaning rule matching according to claim 1, characterized in that: After the memory segment is evenly divided according to the number of buckets for bucketing the matching instances obtained from the preset graph data by the instance bucketizer to obtain memory slots corresponding to the number of buckets, the method further includes: Obtain a matching instance, calculate the bucket identifier to which the matching instance belongs, and trigger the instance data corresponding to the matching instance to be written into the bucket corresponding to the bucket identifier; Increasing the memory usage count of the bucket, and detecting whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instance; If it is detected that the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instance, a write offset is determined according to the memory occupancy count, and the instance data corresponding to the matching instance is written into the memory slot corresponding to the bucket according to the write offset.

3. The optimization control method for graph cleaning rule matching according to claim 2, characterized in that: After detecting whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instance, the method further includes: If it is detected that the memory slot corresponding to the bucket is not sufficient to accommodate the instance data corresponding to the matching instance, it is determined that the memory slot is full, and the data in the memory slot is appended to the file corresponding to the disk; After the writing pointer of the memory slot is updated after the appending is completed, the method further includes: According to the write pointer, the instance data corresponding to the matching instance is written into the memory slot corresponding to the bucket.

4. The optimization control method for graph cleaning rule matching according to claim 2, characterized in that: The increasing the memory usage count of the bucket includes: Obtaining a current occupancy count of the bucket and an instance size of instance data corresponding to the matching instance; Based on the atomic variable, using the instance size, the current occupancy count is increased to obtain an updated occupancy count, and according to the updated occupancy count, the detection of whether the memory slot corresponding to the bucket is sufficient to accommodate the instance data corresponding to the matching instance is performed, and the current occupancy count is returned; The determining the write offset according to the memory occupancy count includes: The write offset is determined based on the returned current occupancy count.

5. The optimization control method for graph cleaning rule matching according to claim 4, characterized in that: After appending the data in the memory slot to the file corresponding to the disk, the method further includes: Reset the current occupancy count formed by the atomic variable to 0, and restart the write thread blocked by the condition variable; After the writing pointer of the memory slot is updated after the appending is completed, the method further includes: Based on the write thread, and according to the write pointer, restart the bucket write operation.

6. The optimization control method for graph cleaning rule matching according to any one of claims 1 to 4, characterized in that: If all matching instances are single-beam structure tables, the single-beam structure table includes a central node and at least one leaf node, the central node and each leaf node correspond to an attribute respectively, and each attribute corresponds to a data value under any instance data, after writing the instance data into the memory slot corresponding to the bucket, it also includes: Compressing the matching instances in the memory slots of the buckets, and taking the union of the attributes of all the single-bundle structure tables in all the matching instances as the header of the conditional compaction table; All matching instances with the same central node as the root are stored in the same row of the condition compact table to obtain a compressed condition compact table.

7. The optimization control device for graph cleaning rule matching according to claim 6, characterized in that: After obtaining the compressed conditional compact table, the method further includes: Acquire the attributes of each matching instance and the predicate dependency conditions that satisfy the corresponding attributes, and form a conditional annotation of the corresponding matching instance according to the attributes and the predicate dependency conditions; The conditional annotations of all matching instances are matched to the compressed conditional compact table, so that the corresponding matching instances are reconstructed according to the conditional annotations and the compressed conditional compact table.

8. An optimization control device for graph cleaning rule matching, characterized in that: include: An initialization module, used to initialize a continuous memory area as a memory pool, and directly select a memory segment from the memory pool when a graph cleaning rule matching task is submitted; A memory slot construction module, used to evenly divide the memory segment according to the number of buckets into which the instance bucketizer buckets the matching instances obtained from the preset graph data, and obtain memory slots corresponding to the number of buckets; A data writing module is used for writing the instance data into a memory slot corresponding to any bucket after the instance data is acquired in the bucket; A repeated persistence module is used to append the data in any memory slot to the file corresponding to the disk after the memory slot is full, and to update the write pointer of the memory slot after the appending is completed, and the write pointer is used to indicate that subsequent data is directly written; The entire persistence module is used to write all the remaining data in all memory slots into the corresponding files on the disk after all the data has been bucketed.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the optimization control method for graph cleaning rule matching as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the optimization control method for graph cleaning rule matching as described in any one of claims 1 to 7 is implemented.