Automatic approximate refinement of rule set from annotated data set
By using a boundary structured rule representation method, the input dataset is transformed into a K-dimensional sub-region and the boundary is extended to generate a rule set. This solves the problems of overfitting, underfitting, and mislabeling in existing mapping techniques, and achieves efficient and reliable approximate mapping and incremental refinement.
Patent Information
- Application Number
- CN202480032109.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-04
- Filing Date
- 2024-05-02
- Publication Date
- 2026-01-13
AI Technical Summary
Existing technologies suffer from overfitting, underfitting, nonlinear decision boundaries, and performance degradation due to increased dimensionality when generating mappings from multiple input values to output values. Furthermore, traditional methods struggle to handle mislabeling errors and incremental refinement of mappings.
The Boundary Structured Rules (BSR) notation is adopted. By converting the input dataset into K-dimensional sub-regions and expanding these regions according to the approximate bounds, a set of rules is generated to achieve approximate mapping, providing reliable approximate bounds and incremental refinement.
It achieves the generation of efficient and reliable approximate mappings while reducing computing and storage resources, adapting to the uncertainty and error of input data, supporting incremental refinement, and reducing design and runtime.
Smart Images

Figure CN121336216A_ABST
Abstract
Description
[0001] Cross-references to other applications This application claims priority to U.S. Provisional Patent Application No. 63 / 463,986, filed May 4, 2023, entitled “AUTOMATIC APPROXIMATING REFINEMENT OF RULESET FROM LABELLED DATA SETS”, which is incorporated herein by reference for all purposes. Background Technology
[0002] Various applications require mappings from multiple input values to output values, such as in classification and / or decision making. Providing such mappings would be an improvement, in part using fewer computational resources, including fewer processing resources, fewer storage resources, and / or fewer network resources, and / or reducing design time and / or reducing the runtime of the mapping application. Attached Figure Description
[0003] Various embodiments of the present invention are disclosed in the following detailed description and accompanying drawings.
[0004] Figure 1 This is a functional diagram illustrating a programming computer / server system for mapping, based on some embodiments.
[0005] Figure 2 This is a diagram illustrating an example of a conventional mapping.
[0006] Figure 3 This is a diagram illustrating an example of a boundary structured representation.
[0007] Figure 4 This is a diagram illustrating an example of transforming labeled points into a structured representation of point boundaries.
[0008] Figure 5 This is a diagram illustrating an example of conversion to an expanded labeled area.
[0009] Figure 6A and Figure 6B This is a flowchart illustrating an example of a process for performing basic automatic approximate refinement of a rule set based on a labeled dataset.
[0010] Figure 7 It is a diagram illustrating the basic automatic generation and / or approximate refinement of the rule set based on the labeled dataset.
[0011] Figure 8-11 The figure illustrates the basic automatic generation and / or approximate refinement of the rule set based on the labeled dataset using unidentified subregions.
[0012] Figure 12A and12B This is a flowchart illustrating an example of an automatic approximate refinement process for a rule set based on a labeled dataset, performed using bucketed data.
[0013] Figure 13A and 13B This is a flowchart illustrating an example of a process for pre-grading and automatically approximating and refining a rule set based on a labeled dataset.
[0014] Figure 14 This is a flowchart illustrating an example of an automatic approximate refinement process for a rule set based on a labeled dataset. Detailed Implementation
[0015] This invention can be implemented in numerous ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by memory coupled to a processor. In this specification, these embodiments or any other form of the invention may be referred to as technology. Generally, within the scope of this invention, the order of steps of the disclosed processes may be varied. Unless otherwise stated, components such as processors or memory described as being configured to perform tasks may be implemented as general components temporarily configured to perform tasks at a given time, or manufactured as specific components to perform tasks. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data such as computer program instructions.
[0016] The following provides a detailed description of one or more embodiments of the present invention, along with accompanying drawings illustrating the principles of the invention. The invention has been described in conjunction with such embodiments, but is not limited to any particular embodiment. The scope of the invention is defined only by the claims, and the invention covers numerous alternatives, modifications, and equivalents. Numerous specific details are set forth in the following description to provide a thorough understanding of the invention. These details are provided for illustrative purposes, and the invention can be practiced according to the claims without requiring some or all of these specific details. For clarity, technical materials known in the art related to the invention have not been described in detail so as not to unnecessarily obscure the invention.
[0017] It discloses mappings from multiple input values to output values, such as in classification and decision making.
[0018] Figure 1 This is a functional diagram illustrating a programming computer / server system for mapping, based on some embodiments. As shown, Figure 1Functional diagrams of general-purpose computer systems programmed to provide mappings according to some embodiments are provided. As will be apparent, other computer system architectures and configurations can be used for mappings.
[0019] Computer system 100 includes various subsystems as described below, including at least one microprocessor subsystem, also referred to as a processor or central processing unit (“CPU”) 102. For example, processor 102 may be implemented by a single-chip processor or a multi-core and / or processor. In some embodiments, processor 102 is a general-purpose digital processor that controls the operation of computer system 100. Using instructions retrieved from memory 110, processor 102 controls the reception and manipulation of input data, as well as the output and display of data on an output device, such as a display and graphics processing unit (GPU) 118.
[0020] Processor 102 is bidirectionally coupled to memory 110, which may include a first main storage device (typically random access memory (“RAM”)) and a second main storage area (typically read-only memory (“ROM”)). As is well known in the art, the main storage device can be used as general storage and scratch-pad memory, and can also be used to store input data and processed data. In addition to other data and instructions used for processes operating on processor 102, the main storage device may also store programming instructions and data in the form of data objects and text objects. Again, as is well known in the art, the main storage device typically includes basic operating instructions, program code, data, and objects, such as programming instructions, used by processor 102 to perform its functions. For example, as described below, depending on whether data access requires bidirectional or unidirectional access, main storage device 110 may include any suitable computer-readable storage medium. For example, processor 102 may also directly and very quickly retrieve frequently needed data and store it in cache memory (not shown). Processor 102 may also include a coprocessor (not shown) as an auxiliary processing component to assist processor and / or memory 110.
[0021] Removable mass storage device 112 provides additional data storage capacity to computer system 100 and is coupled to processor 102 bidirectionally (read / write) or unidirectionally (read-only). For example, storage device 112 may also include computer-readable media such as flash memory, portable mass storage devices, holographic storage devices, magnetic devices, magneto-optical devices, optical devices, and other storage devices. For example, fixed mass storage device 120 may also provide additional data storage capacity. An example of mass storage device 120 is an eMMC or microSD device. In one embodiment, mass storage device 120 is a solid-state drive connected by bus 114. Mass storage devices 112 and 120 generally store additional programming instructions, data, and the like that are not currently actively used by processor 102. It will be appreciated that, if needed, the information held in mass storage devices 112 and 120 can be used as virtual memory in a standard manner as part of main storage device 110 (e.g., RAM).
[0022] In addition to providing processor 102 with access to the storage subsystem, bus 114 can also be used to provide access to other subsystems and devices. As shown, these may include display monitor 118, communication interface 116, touch (or physical) keyboard 104, and one or more auxiliary input / output devices 106, including audio interfaces, sound cards, microphones, audio ports, audio input devices, speakers, touch (or pointing) devices, and / or other subsystems as needed. Besides a touchscreen, auxiliary devices 106 may be a mouse, stylus, trackball, or tablet, and are useful for interacting with a graphical user interface.
[0023] Communication interface 116 allows processor 102 to be coupled to another computer, computer network, or telecommunications network using the illustrated network connection. For example, through communication interface 116, processor 102 can receive information (e.g., data objects or program instructions) from another network, or output information to another network during the execution of method / process steps. Information typically represented as a sequence of instructions to be executed on the processor can be received from and output to another network. Interface cards or similar devices, along with appropriate software implemented by processor 102 (e.g., executed / implemented on processor 102), can be used to connect computer system 100 to an external network and transfer data according to standard protocols. For example, various process embodiments disclosed herein can be executed on processor 102, or can be executed across a network such as the Internet, intranet, or local area network by a remote processor in conjunction with a portion of the shared processing. Throughout this specification, "network" refers to any interconnection between computer components, including the Internet, Bluetooth, WiFi, 3G, 4G, 4G LTE, GSM, Ethernet, intranet, local area network ("LAN"), home area network ("HAN"), serial connection, parallel connection, wide area network ("WAN"), Fibre Channel, PCI / PCI-X, AGP, VLbus, PCI Express, Expresscard, Infiniband, ACCESS.bus, wireless LAN, HomePNA, fiber optic, G.hn, infrared network, satellite network, microwave network, cellular network, virtual private network ("VPN"), universal serial bus ("USB"), FireWire, serial ATA, I-Wire, UNI / O, or any form that connects homogeneous and / or heterogeneous systems and / or groups of systems together. Additional mass storage devices (not shown) may also be connected to processor 102 via communication interface 116.
[0024] An auxiliary I / O device interface (not shown) can be used in conjunction with computer system 100. The auxiliary I / O device interface may include general and custom interfaces that allow processor 102 to send and more typically receive data from other devices such as microphones, touch-sensitive displays, transducer card readers, tape readers, voice or handwriting recognition devices, biometric readers, cameras, portable mass storage devices, and other computer components.
[0025] Furthermore, the various embodiments disclosed herein also relate to computer storage products having a computer-readable medium, which includes program code for performing various computer-implemented operations. A computer-readable medium is any data storage device capable of storing data that can subsequently be read by a computer system. Examples of computer-readable media include, but are not limited to, all of the media mentioned above: flash memory media, such as NAND flash, eMMC, SD, compact flash; magnetic media, such as hard disks, floppy disks, and magnetic tapes; optical media, such as CD-ROM discs; magneto-optical media, such as optical discs; and specially configured hardware devices, such as application-specific integrated circuits (“ASICs”), programmable logic devices (“PLDs”), and ROM and RAM devices. Examples of program code include both: machine code generated by a compiler, for example, or files containing higher-level code, such as scripts that can be executed using an interpreter.
[0026] Figure 1 The computer / server system shown is merely an example of a computer system suitable for use with the various embodiments disclosed herein. Other computer systems suitable for such use may include additional or fewer subsystems. Furthermore, bus 114 is an illustration of any interconnection scheme for linking subsystems. Other computer architectures with different subsystem configurations may also be utilized.
[0027] Determining a mapping with multiple input values can be complex. For example, consider implementing a predictive model of a chemical process as a time-stepped computer simulation. The state of the process at the next state depends on a large number of parameter values corresponding to its current state. If there are... K If there are such parameters, then the mapping is to K Each parameter is mapped to M A mapping of output values, typically partially corresponding to the expected values of these parameters in the next time step and the simulated output.
[0028] In many cases, there may not be a known closed-form formula to calculate the output value from the input values. Closed-form formulas normally rely on the mapping being "continuous" in a mathematical sense. However, in many practical applications, singularities that violate this continuity requirement may exist. Furthermore, many assumptions made by closed-form formulas are not strictly true in real / practical systems. For example, a formula might assume that friction is constant or linear, an assumption that may not be true.
[0029] In many applications, it is feasible to generate labeled datasets from entities in the domain of interest (such as aircraft, seacraft, landcraft, or similar engineered systems) or their intended environment by collecting telemetry data from sensors connected to the system of interest. A labeled dataset consists of a set of records, each typically indicating a specific value of a relevant parameter at a specified time, and labels characterizing or classifying these parameter values. A dataset can be generated in part by equipping the actual system with instruments and collecting data based on operating the system under different circumstances or conditions. It can also be generated in part by collecting the same data based on detailed simulations of the target system. A labeled dataset can also be generated in part by specifying one or more of the input attributes as generating labels and treating the remaining attributes as input.
[0030] In these applications, various conventional techniques can be used to generate a hyperplane as a decision boundary that divides the entire space into two regions, one with an associated label and the other with one or more remaining labels. This process is repeated for the latter region as needed to further subdivide it until a single label is associated with each region. However, this conventional approach suffers from overfitting, underfitting, nonlinear decision boundaries, and / or increasing cost / decreasing performance with increasing dimensionality. This approach finds a hyperplane that maximizes the distance between dissimilarly labeled data points by adjusting weights, thus "fitting" the hyperplane to the data. This "fit" might mean that when called with input values contained in records having those data values, the mapping outputs the correct label. However, there is no guarantee of label correctness for input values not in the training input data. Because conventional techniques use higher-order continuous functions for this mapping, as well as kernel functions, it is infeasible to even approximately predict what the fitted mapping will output for values outside the training set.
[0031] Traditional machine learning / artificial intelligence. Figure 2 This is a diagram illustrating an example of a traditional mapping. Figure 2 In this context, a set of labeled points and / or classification points (202), (204), (206), (208), (210), (212), (214), and (216) is depicted based on the input dataset and / or the set of input records. There are no restrictions. Figure 2 The set of annotation points / input records is shown in a two-dimensional space with two parameters, "X" and "Y".
[0032] Each marker has its own associated label. Figure 2The labels described as "diagonal shading" (202), (204), (206), and (208), or "vertical shading" (210), (212), (214), and (216). Figure 2 In the example shown, the simple basic natural / physical / engineering mapping is that any point above the horizontal / physical classifier (220) is mapped to a "diagonal shading", and any point below the horizontal line (220) is mapped to a "vertical shading".
[0033] In traditional techniques, such as the "machine learning problem," training is performed on a set of input records (202), (204), (206), (208), (210), (212), (214), and (216) to infer the labels of future input points. For example, in machine learning, a complex collection of neurons can be adjusted to maintain equilibrium until consistency with the input data points (202), (204), (206), (208), (210), (212), (214), and (216) is achieved. Figure 2 Let the machine learning classifier (230) be represented in the middle, such that future input points above the machine learning classifier (230) are mapped as “diagonal shading” and future input points below the machine learning classifier (230) are mapped as “vertical shading”.
[0034] like Figure 2 What can be seen in Figure 2 Misclassification or mislabeling errors occur in the regions (242), (244), (246), and (248) depicting the physical classifier (220) and the machine learning classifier (230). That is, if a future input point is within such a region (e.g., region (242)), the machine learning technique will label it as a "vertical shaded" point, even if the underlying natural / physical / engineering property should be "diagonal shaded." Therefore, as... Figure 2 As shown, incorrect labels may be generated by using conventional fitting methods (such as machine learning artificial intelligence (AI) methods).
[0035] Determining the boundaries of mislabeling errors a priori using mappings generated using traditional machine learning / AI can be challenging or impossible. Furthermore, many traditional techniques require a complete regeneration of the mapping when additional labeled input data is available, rather than allowing for incremental refinement.
[0036] In many applications, absolute precision may not be necessary. For example, when using input values of features on manufacturing equipment to classify whether it is about to fail or fail, a small number of false positives may be acceptable, provided that the positive indications of the generated rules allow technicians to make further evaluations. Technicians may only focus on those pieces of equipment classified as potentially failing. This is a substantial improvement compared to technicians having to periodically inspect all pieces of equipment. Furthermore, the set of input datasets / input records may contain some degree of error due to the inaccuracy of sensor readings. Additionally, if the size of the dataset does not fully cover the total input space of possible value combinations (which is normally impractical), it may be impossible to determine a completely accurate mapping between labels and data. Therefore, as long as the approximation error is bounded, generating rules that provide an approximately accurate mapping without absolute precision remains useful.
[0037] A method is disclosed for refining from the training input dataset / set of input records. K An automated method for a reliable approximate mapping from input to output. It can provide reliable bounds on the degree of approximation and / or inaccuracy of the mapping relative to the input training set. As mentioned in this paper, the terminology... Refine This includes full generation, because the disclosed techniques can start with a zero (null) set of maps / rules and refine it to include one or more maps / rules generated from a set of input records.
[0038] Boundary structured representation. Figure 3 This is a diagram illustrating an example of a boundary structured representation (BSR). For example... Figure 3 As shown in the example, the input dataset points (202), (204), (206), (208), (210), (212), (214), and (216) and the physical mapping (220) are similar to Figure 2 The content described in it.
[0039] The article mentioned area It is a closed geometric volume; for example, a three-dimensional region is a closed space. As mentioned in this article, N 3D space Superregion All that is enclosed by a specific boundary N A set of labeled points. In the case of degradation, the superregion consists of a single... N Composed of dimensional points. Prefix "super" (" hyper The '' sign is used to indicate that a space and / or region can be of any dimension, not just 2 or 3. As mentioned in this article, boundary It is specified by one or more boundary structuring rules.
[0040] In one embodiment, the Boundary Structured Rule Set (BSRS) is refined by converting the rules in the BSRS into... K Subregions in 3D space, and then read in according to each record having K The training set consists of input records with corresponding attributes and labels. This is achieved by transforming each input record into... K Subregions are expanded based on application-specific approximate boundaries, merging existing regions into existing regions by adjusting these existing regions and / or forming new regions, and outputting the resulting set of each labeled subregion as a refined rule set. In the special case where no pre-existing rule set exists, the BSRS to be refined can be defined as a null set or a single rule that maps the entire input domain to default labels.
[0041] This can be viewed as interpreting each labeled input record as in K Rules for defining labeled sub-regions in dimensional space are then used to add or merge these "input" sub-regions into the expanded space. K Within a subregion, a rule set is output that comprises a smaller number of rules compared to the number of records in the input dataset, while providing coverage beyond the data points specified in the labeled input data. In a sense, it extends the coverage of one or more bounding geometries of the input rules, consistent with the remainder of the input record set, while merging multiple extended subregions when the merged subregions can be represented by one or more bounding structured rules in the used bounding structured rule system and the merged subregions conform to approximate bounds. As mentioned herein, the terminology used... subregion Instead area This indicates that, from the perspective of the final output, the defined region is not necessarily the final version.
[0042] As mentioned in this article, Boundary structuring A rule is a rule with an antecedent, which specifies the conditions under which a rule is set in order to be ... K Define and enclose sub-regions in the dimensional input space Boundary geometry And the Boundary structuring The consequent of a rule similarly specifies the condition in which the condition is met. M Define and enclose sub-regions in the dimensional output space Boundary geometryIn fact, a boundary structuring rule instructs each data point within the boundary geometry of its predecessor / subregion to be mapped to a point in the boundary geometry of its successor. Boundary geometry can indicate a single point or a multi-point subregion. A boundary structuring rule can be matched with data points in a variety of disjoint subregions by its predecessor, corresponding to the disjunction of the boundary specification of each of these subregions. As mentioned in this paper, extending a subregion means modifying the boundary of the subregion to enclose additional points. Extending a boundary structuring rule means modifying a rule to specify an extended corresponding boundary. An extended boundary structuring rule set means extending the rules in the rule set or adding new boundary structuring rules to the rule set, or both.
[0043] The rule set of boundary-structured rules partially defines the mapping from the input parameter space to the output result space of the mapping. In particular, when the mapping is invoked with a given set of parameter values, the rules are evaluated, and for each rule whose antecedent evaluates to true (i.e., the parameter value is contained in the sub-region defined by the antecedent), the consequent contributes directly or indirectly to the output result.
[0044] To simplify the description without loss of generality or limitation, as described below, labels are described as corresponding to a single attribute. For example, given the values of other attributes in the input record, a label could indicate an expected temperature change. In a classification dataset of fault conditions, a label value could indicate the category of risk associated with the fault. This single label value can be assumed without loss of generality for several reasons. First, multiple individual values can be presented as... M A tuple or composite value is a combination of these individual values. For example, a label might specify a record containing all three of temperature, pressure, and humidity. Approximate bounds can be calculated based on the sub-attributes of these individuals. Alternatively, the disclosed techniques can be applied to separately... M Each of the output values generates an approximate rule-based mapping, thus providing... M Mapping of output values. Then, M Individual rule sets can be combined to produce a rule set that maps to these rules in each rule evaluation cycle. M A separate attribute. As a separate attribute M There are no conflicts between the individual results.
[0045] Boundary structuring rules are represented in various ways to specify boundary geometry. For example, range structuring rules are boundary structuring rules where the boundary geometry is an axis-aligned hyperrectangle or orthotope. Therefore, each dimension has a clause, and each clause indicates the minimum and maximum values of the dimension / attribute associated with that dimension.
[0046] For example, This is a boundary structuring rule, where `stoveReady` is a data point (enumerated value or category) in the output / result space, and it is mapped by data points in subregions where `stoveFaults`, `pilotLight`, and `gasFlow` are all within the constraints specified in the antecedent. This boundary structuring rule can be transformed into a labeled region of space with dimensions `stoveFaults`, `pilotLight`, and `gasFlow`, containing all points in subregion [0,3) with `stoveFaults` dimension values, in range [10,20) with `pilotLight` values, and in range [4,9) with `gasFlow` values. In the notation used in this paper, the range specified as [threshold 0, threshold 1) indicates values greater than or equal to threshold 0 and strictly less than threshold 1. Therefore, if the attribute is an enumeration of possible states, and 1 is a numeric specification for normal operation, then the range [1,2) is a range specification for the value 1. Generally, in its antecedent, it has... K The scope of each clause is structured according to rules, with each clause associated with a different dimension, defining what is... K 3D hyperrectangle K Dimensional boundary geometry.
[0047] Continuing with the example above, the reverse rule applies: It is also a boundary structuring rule, where stoveReady is a data point in the input parameter space, such as an enumeration value or category, and is mapped to a sub-region, where stoveFaults, pilotLight, and gasFlow are all within the constraints specified in the consequent.
[0048] In one embodiment, the range structure rules are transformed and / or converted into annotations. K The subregion, referred to in this paper as Hyperrectangle Furthermore, the labeled axis-aligned hyperrectangles can be transformed and / or converted into range structured rules. Similarly, range structured rules can correspond to diverse subregions by disjunctions whose antecedents are conjunctive clauses (a conjunctive clause for each subregion).
[0049] In one embodiment, an alternative representation of the boundary structuring rules is that the boundary geometry is represented as a vector representation of a convex polyhedron, i.e. KA set of extrema. This allows for a more flexible definition of the bounding geometry, since each hyperrectangle is a convex polyhedron, not the other way around. However, this can also make determining whether a data point is contained within a given convex polyhedron and / or for related operations such as intersection more expensive. Using a vector representation of a convex polyhedron, the boundary structuring rules equivalent to those for a single data point are achieved by repeating the value of each dimension D in the dimension D component of each extremum point.
[0050] Besides specifying the vector representation of a hyperrectangle or convex polyhedron, there are unrestricted alternatives to specifying it. K Bounded geometry in 3D space. For clarity, as described herein, a hyperrectangle may be used, for example, but other representations may be used without restriction. Convex bounded geometry can generally determine the overlap, intersection, and extension of geometry more efficiently, but non-convex geometry may also be used without restriction alternatively.
[0051] The description in this article Boundary Structured Rule System It is a system with a rule evaluator that can determine whether input data points fall within the bounded geometry specified in the rule antecedent, and whether the two bounded geometry specifications intersect. In other words, Boundary structuring A rule system restricts the specification of the boundary, and thus the bounding geometry, enabling it to efficiently evaluate rule antecedents. For a given bounding structured rule system, these restrictions limit what the bounding geometry can be represented. For example, a range structured rule system could be a bounding structured rule system that represents axis-aligned hyperrectangles only in its antecedent and consequent.
[0052] Figure 4 This is a diagram illustrating an example of converting labeled points to a structured representation of point boundaries. For example... Figure 4 As shown in the example, the two-dimensional XY axis (where K =2) Similar to Figure 2 The axis depicted in the text.
[0053] From such as Figure 4 The system shown has input data points of non-trivial size. K In a subregion, such that it has the same or similar consequent for most points within that subregion. For the simple two-dimensional case, in Figure 4 The diagram illustrates this transformation / conversion to a subregion. Here, the asterisk (402) indicates the actual data point, and the surrounding rectangle (404) indicates the region surrounding the data point, which has the same label due to this local statistical smoothness or local continuity property. Figure 4 The remaining areas depicted are labeled "L1".
[0054] The set of input records is transformed into labeled sub-regions that serve as boundary structuring rules. This set of input records can be assumed or required to be structured as a sequence of records, where each record is a sequence of label values and values, with each value corresponding to a specific label. property and / or feature The attributes or features are referred to herein as input parameters. It is further assumed that attribute information exists for each attribute, providing indications of its minimum and maximum values. This also includes approximate limits as indicators of accuracy (such as the required accuracy). For example, a temperature sensor can provide a 16-bit data value as its reading. However, the full range might be limited to only -40 degrees to 120 degrees, with an accuracy of + / - 0.10 degrees. The 16-bit value received from the sensor can indicate the temperature in hundreds of degrees relative to -40 degrees. The temperature value can then be converted into a range [value -0.10, value +0.10]. Depending on the accuracy characteristics, the actual range used in the dimensions of this resulting sub-region may be extended.
[0055] Therefore, it is necessary for the input record of a single data point to be transformed into K A labeled subregion in 3D space. K The dimensional subregion is nontrivial, meaning it has a non-zero volume because of the permissible approximate bounds in each dimension, such as... Figure 4 As illustrated in the diagram. Figure 4 In this model, by extending in two dimensions, two-dimensional data points are transformed into two-dimensional labeled sub-regions of a hyperrectangular structured rule system. The approximate bound is given by the maximum increment (delta) of the dimensions, such as... Figure 4 As shown in the diagram, the X and Y dimensions of the data points are transformed into a range from [x-deltaX, X+deltaX), as... Figure 4 The element (406) in the table is shown. The Y dimension of the data points is transformed into a range from [y-deltaY, Y+deltaY), as shown in the table. Figure 4 The element (408) in the middle is shown.
[0056] An approximation limit can represent both the potential inaccuracy of the input data and / or the range across which the difference in values is insignificant under normal circumstances. For example, regarding the latter, ambient temperature can be measured to a value of 0.10 degrees, but for the system being measured, an ambient temperature value within a range of 5 degrees above or below is insignificant under normal circumstances. Therefore, an approximation limit can be specified as adding or subtracting 5 degrees.
[0057] In one embodiment, approximate bounds are provided in various different forms. One form is a function that returns approximate bounds above and below specified values based on the entire input record.
[0058] It is assumed and required that a value exists in the input record for each attribute. This value can indicate an "ignore" value, indicating that the label is true for any possible corresponding attribute value. If the value is otherwise unknown, preprocessing can fill in possible values or default values.
[0059] Through K Defining subregions in dimensional space allows boundary structuring rules to be transformed into labeled subregions whose boundaries correspond to the boundaries specified by the rule, and whose labels match the labels of the rule consequents.
[0060] The labeled sub-region representation can be stored separately from its rule representation. However, in such embodiments, there are automatic methods to convert and / or transform the labeled sub-region representation into the rule representation, and vice versa, without significant errors. In alternative embodiments, the labeled sub-region representation is merely an internal representation (such as a range list and labels), which can be considered as labeled sub-regions or boundary structured rules.
[0061] Local statistical smoothness. For many applications, especially those using engineered and / or physical systems, fundamental natural mechanisms provide a mathematically defined... Piecewise continuity Behavior and Smoothness Therefore, for a given mapping and data points (DPs) defined by one or more input parameter values, K Many nearby data points in the dimensional space are mapped to approximately the same output; that is, the output is the data point adjacent to the data point mapped by DP. This is referred to in this paper as... Local statistical smoothness This is because most, but not all, of the neighboring data points may be similar to that data point in terms of label.
[0062] In one embodiment, two subregions with the same or nearly the same labels may be merged into one subregion if they match or overlap in one dimension and are approximately matched in other dimensions. This is partly because otherwise, the system might exhibit discontinuous or highly dynamic behavior between these data points, challenging the simplification of local statistical smoothness. This is analogous in a sense to stating that behavior is approximately transitive; that is, if input I1 is “close” to input I2, and input I2 is “close” to I3, then I1 is “close” to I3, as mentioned herein. nearA is close to B when A and B are within the approximate bounds of the real system being represented. However, two subregions with different labels may be adjacent to each other to indicate discontinuities across their adjacent boundaries, thus potentially reducing this extension.
[0063] Figure 5 This is a diagram illustrating an example of transformation to an expanded labeled area. For example... Figure 5 As shown in the example, the two-dimensional XY axis (where K =2) Similar to Figure 2 The axis depicted in [the text]. Figure 5 The diagram illustrates transitivity and discontinuity. Each of the existing and / or extended labeled subregions ELR0 (502), ELR1 (504), ELR2 (506), and ELR3 (508) is a subregion subsumed and / or merged based on two or three separate data points, and is labeled “diagonal shading” for ELR2 (506) or “vertical shading” for ELR0 (502), ELR1 (504), and ELR3 (508). The remainder of the region is labeled “L1”. The different labels for ELR2 (506) indicate inconsistencies and / or discontinuities (510) between the ELR1 (504) and ELR2 (506) subregions along the X-axis. Figure 5 As described in the document, ELR3 (508) is too far from ELR0 (502) and ELR1 (504) to be merged with those subregions.
[0064] In one embodiment, the rule set can cover the input domain by providing boundary-structured rules for each distinct “piece” or sub-region corresponding to the segmented continuous behavior, that is, each sub-region exhibits relatively “smooth” behavior, as described herein. smooth Refers to continuous and approximately identical behavior. There may be multiple sub-regions with the same label, as shown using ELR3 (508) and ELR0 (502).
[0065] Approximately correct mappings within certain approximate limits are sufficient for many practical applications. For example, in control applications, the input state changes continuously, thus requiring periodic control actions. If, due to this approximation, the control action is undercorrected at one time step, it can be compensated for by further correction at subsequent time steps. By restricting the merging of sub-regions to these approximate limits on both data points and labels, a rule set of boundary structuring rules can be used to achieve mappings that satisfy the approximation requirements of these applications.
[0066] Approximate bounds for each attribute. In one embodiment, there is an approximate bound for each attribute. It is recognized that the relevant error level and application-level sensitivity are typically specific to the source of the data and the nature or semantics of the attribute. Therefore, if the difference between the values of each attribute A in data points DP0 and DP1 is less than the corresponding approximate bound for A, then data point DP0 is considered as described herein. Neighbor For data point DP1, the approximate bound can be chosen based on knowledge of the application domain (especially the dynamics of the system over time). It can also be calculated based on the values of other attribute values.
[0067] This method contrasts with the traditional approach of calculating Euclidean distance based on differences calculated across dimensions. For example, if one attribute is color and the other is temperature, the distance metric calculated based on these two attributes lacks a clear basis or semantics.
[0068] There are also approximate limits regarding label values. If an input record maps to an existing subregion—that is, its corresponding subregion intersects with that existing subregion—and the difference between its label and the label of that subregion is less than the label approximate limit, then a conflict can be resolved by considering the input record as having the label of the existing subregion. For example, if the input label is "moderately heavy" and the existing label is "slightly heavy," and these labels are considered sufficiently close, then if the input label maps to a subregion labeled "slightly heavy," especially if the input label is an isolated point within a subregion labeled "slightly heavy," then the input label might alternatively be considered as being designated "slightly heavy." If the difference between the label and the label of the subregion to which the input record data point is mapped exceeds the approximate limit, then a conflict is considered to exist. conflict They also demanded a resolution to the conflict.
[0069] Basic automatic generation method. Figure 6A and 6B This is a flowchart illustrating an embodiment of a process for performing basic automatic approximate refinement of a rule set based on a labeled dataset. In one embodiment, by... Figure 1 System processing Figure 6A and 6B The process.
[0070] In step (602), the existing Boundary Structured Rule Set (BSRS) and associated attribute information are input, such as the minimum and maximum values of each attribute and the approximate limits of the attributes. An internal representation corresponding to the labeled sub-region is generated, such as... Figure 4 and Figure 5The following is a simplified description. In step (604), the set of input records is accessed. In step (606), an attempt is made to read the "next input record", which is initially the first input record, and a loop is established over all input records between steps (608) and (612) to (622).
[0071] Next, in step (608), a check is performed to determine if a valid next input record exists. If not, in step (610), the set of labeled sub-regions is output as a refined version of the BSRS; otherwise, control is transferred to step (612). In optional step (612), the input is generated using the record's label and the position of the data point specified by the value in the input record. Label sub-region (ILR), this input Label sub-region (ILR) is extended by approximate bounds for each dimension. This labeled subregion is represented as, or can be accurately transformed / converted into, boundary-structured rules in a rule system. In alternative embodiments, such as Figure 6B The step (612) shown simply represents the input ILR by the data points themselves, without extending the input ILR by approximate bounds for each dimension.
[0072] In step (614), the input ILR is "enclosed" by the existing labeled subregion (ELR). As mentioned herein, subregion A is enclosed by subregion B if: (1) subregion A is contained within subregion B or an extended subregion B, and is extended in accordance with approximate bounds on dimensions / attributes; and (2) the labels on subregion A are within approximate bounds of the labels on subregion B. Package Photo .
[0073] In step (614), if the ILR is covered by the ELR, then in step (616), the ILR is merged into the ELR; otherwise, control is transferred to step (618). In step (618), if the ILR conflicts with the ELR due to significantly different labels, then the conflict is resolved in step (620); otherwise, control is transferred to step (622). In step (620), if the conflict is resolved in the new region, control is transferred to step (622); otherwise, control is transferred to step (624). In step (622), a subregion is added to the set of existing subregions, wherein the subregion corresponds to the ILR and its label corresponds to the ILR label. In step (624), the ILR is ignored and / or eliminated.
[0074] review Figure 6A and 6B The process, given input data points and their associated ILR regions: 1. If it has the same or nearly the same label as the ELR and is contained within the ELR, then it is encompassed by the existing label subregion of the ELR, or the ELR may be extended to include it without exceeding the specified approximate limit; 2. If it overlaps with an ELR sub-region with a different label, a conflict occurs, and the conflict resolution mechanism is invoked to resolve it; or 3. If it is not merged and is not eliminated as a conflict, it is added as a subregion to the set of existing labeled subregions.
[0075] Therefore, informally, if it is the same label and close to it, the new input record expands the existing sub-region, and otherwise creates a new labeled sub-region, unless it is eliminated by the conflict resolution. As mentioned in this article, elimination can include recording an incorrect or questionable input value, or simply discarding the input.
[0076] return Figure 5 This illustration demonstrates the use of a super-rectangular BSR system, through Figure 6A and Figure 6B The process generates extended labeled sub-regions: ELR0 (502) is an ELR with the "vertical shadow" label, generated by merging three adjacent ILRs; ELR1 (504) is also an ELR with the "vertical shadow" label, generated from two distinct adjacent ILRs. However, even though it is adjacent to ELR0, it will not merge with ELR0 because the result will not be a hyperrectangle; ELR2 (506) is another ELR generated from two adjacent ILRs with the "diagonal shading" label, but with different labels, and therefore is not merged with ELR1; ELR3 (508) is a separate ELR with the same "vertical shadow" label as ELR0 and ELR1, generated by two adjacent ILRs, but not merged with ELR0 and ELR1 because it is not adjacent.
[0077] In one embodiment, conflict resolution records a conflicting subregion of the input label when invoked and may mark it as not to be considered further, that is, eliminate and / or ignore it. If the result is not excessively inconsistent with the set of input records, one form of conflict resolution is to reduce the subregion that conflicts with it.
[0078] In one embodiment, the process merges any two existing labeled sub-regions that have the same label, and these two labeled sub-regions are sufficient. NeighborAnd / or close to, and can be represented as BSR in the merge result, and whose merged sub-regions do not produce new conflicts, that is, overlap with another labeled sub-region with a different label. For example, if Figure 5 ELR1 (504) in the data is then passed to another data point at a later point. Figure 5 (not shown in the image) The data point is expanded by adding a sub-region on top of it, so that ELR0 (502) and ELR1 (504) form a super rectangle, and can therefore be merged and represented as a single labeled super rectangle, by Figure 6A and 6B The process of representation discovers this and merges the two sub-regions, thereby generating a merge of the corresponding rules. This checking and merging action for the mergibility of sub-regions can occur during the input of the IRL (602), (604) or as a post-processing step after the process has been completed (610).
[0079] In one embodiment, a conventional collision detection algorithm is used to detect overlap between a new input record sub-region and an existing sub-region. For example, each sub-region may have a bounding geometry that includes all data points within an approximate boundary of that sub-region. Therefore, a conventional collision detection algorithm can be used with these bounding geometries to detect whether a data point specified by the input record is within or sufficiently close to an existing sub-region.
[0080] Figure 7 This is a diagram illustrating the basic automatic generation and / or approximate refinement of the rule set based on the labeled dataset. For example... Figure 7 As shown in the example, the two-dimensional XY axis (where K =2) Similar to Figure 2 The axis depicted in the text Figure 2 A subregion designated by a superrectangle is shown, approximating a continuous curve (702) of the underlying natural / physical / engineered system. In one embodiment, Figure 7 yes Figure 6A and 6B The complete result of the process.
[0081] Enter annotation points and / or enter records in Figure 7 The annotations are depicted as follows: those associated with label L0 are depicted with asterisks; and those associated with label L1 are depicted with diamonds. The labeled sub-regions are... Figure 7 The following content is used to depict the content: those associated with label L0 are depicted with "vertical shading"; and those associated with label L1 are depicted with "diagonal shading".
[0082] Based on local statistical smoothing and approximate bounds Figure 7The complete result has far fewer sub-regions compared to the number of input records, and therefore relatively fewer rules. For example, as Figure 7 As shown, there are eight asterisks but only five L0 marked areas, and 11 diamonds but only five L1 marked areas. Figure 7 The complete results also provide more comprehensive coverage than what a single input record can provide.
[0083] exist Figure 7 In this context, by requiring that the difference between the Y value at the top boundary and the Y value of any data point in the hyperrectangle does not exceed this approximate limit, the approximate limit of the Y value restricts the width of each hyperrectangle. For example, the leftmost L0 rectangle (704) is restricted in width because the L1 data point (706) to its right might otherwise force the top of the L0 rectangle (704) further down compared to the top that the Y approximation bound by the top left L0 data point (708) might allow.
[0084] Note that without approximate boundaries and without transforming individual data points into subregions, there is no basis for a single rule to represent multiple input data points. Therefore, the generated rule set can have as many rules as the number of unique records in the set of input records, and no generated rule can match an input with a combination of values that does not exist in the set of input records being generated.
[0085] Restricted Inheritance and Evolution. In one embodiment, restricted inheritance can be used to generate a boundary structured rule set, as disclosed in U.S. Patent Application No. 18 / 243,631 (Publication No. US / 20240086724A1), filed September 7, 2023, entitled “RESTRICTED INHERITANCE-BASED RULECONFLICT RESOLUTION,” which is incorporated herein by reference for all purposes. Specifically, the generation of the boundary structured rule set can be invoked using an input rule set that specifies the base rule set as well as the set and feature information of the input dataset / input records. Therefore, the refined rule set is generated as a refinement of the base rule set, the derived rule set, and the input rule set.
[0086] The initial subregion is initialized in part based on the rules in the base rule set and the input rule set, where the subregions in the base rules are indicated as base subregions. This base rule set can be generated manually, automatically based on the base set of input records, or a combination of both. For example, the base rule set can be automatically generated by the disclosed method, but with its own manually generated base rule set. As in U.S. Patent Application 18 / 243,631, the set of input records may include additional features or attributes beyond those in the base rule set. In this case, these additional features or attributes are considered "ignorable" in the rules of the rule set, and therefore do not prevent the identification of overlap between derived rules and base rules.
[0087] The generation process is then the same as the previous method, except that when a new candidate BSR subregion overlaps and conflicts with the basic subregion rules, it is instructed to overwrite the basic rules in the overlapping subregion instead of the actually conflicting derived subregion. Input records encompassed by the basic rules are processed as before, that is, such as allowing processing to proceed to the next input record without any further action on that input. The method still checks for conflicts between input record subregions and other subregions associated with the top-level input rule set.
[0088] After the resulting rule set is generated, the rule system evaluates the resulting output rule set as disclosed in U.S. Patent Application No. 18 / 243,631. In particular, when the input matches both the derived rule and the base rule, the consequent of the derived rule is used.
[0089] Use an automatic generation method for unconfirmed sub-regions. In one embodiment, maintain per sub-region. Unconfirmed subregion Therefore, there exists a set of extended labeled subregion pairs (ELRPs), where each ELRP includes... Confirmed Sub-regions and enclosed areas Unconfirmed Subregions. Confirmed subregions are determined using the same method as previously described. Unconfirmed subregions are the maximum extension of the subregion within the constraints of the boundary representation, and do not conflict with confirmed subregions of another ELR.
[0090] In one embodiment, by adapting such as Figure 6A and Figure 6BThe process shown in step (622) maintains unconfirmed subregions to create unconfirmed subregions corresponding to the new / confirmed subregions. This requires starting with a copy of this new / confirmed subregion and expanding it until any further expansion could overlap with one or more other confirmed subregions. Furthermore, it reduces any other unconfirmed subregions so that they do not overlap with the new confirmed subregion.
[0091] Figure 6A and Figure 6B Step (616) in the process is also adapted to reduce any other unconfirmed sub-regions so that they do not overlap with the expanded confirmed sub-regions. Similarly, step (610) is adapted to use unconfirmed sub-regions to output rules. A rule corresponding to an unconfirmed sub-region is a rule in which the antecedent is not significantly contradicted by data points in the set of input records. Therefore, it is not limited to the specified approximate limits. Such a rule, which may have a lower confidence, can be used when the output rule set does not contain rules generated based on confirmed sub-regions that include combinations of the current input values. Therefore, unconfirmed sub-regions allow the output rule set to achieve coverage at the expense of not maintaining approximate limits when the set of input records is not sufficiently "dense," i.e., does not contain enough data points to achieve coverage with specified approximate limits.
[0092] Confidence metric. In one embodiment, for each confirmed / unconfirmed subregion pair, a confidence or accuracy metric or level can be calculated based on the characteristics of the confirmed and unconfirmed subregion pair, such as the maximum distance between the boundaries of the confirmed and unconfirmed subregions, the volume of the unconfirmed subregion relative to the confirmed subregion, the number of data points on the pair or even its adjacent pairs, and other potential metrics. The application can use this as an indication of whether the approximation of the unconfirmed subregions is accurate enough, reporting which subregions in the dataset are not adequately covered. For example, each generated rule can have a confidence level indicated as a function of this metric. If the confidence is too low, rule set generation can be rerun with an expanded dataset, specifically with data from the under-covered regions. Alternatively, additional rules can be provided for these subregions by manually generating additional rules based on application knowledge.
[0093] In one embodiment with a hyperrectangular BSR representation, the new unidentified subregion can initially be identical to the enclosed unidentified subregion. Then, an attribute is selected to divide the new subregion and the enclosed subregion's attribute range. This recognizes that if one subregion does not overlap with another in one dimension, it does not overlap at all.
[0094] illustrate. Figure 8-11This is a diagram illustrating the basic automatic generation and / or approximate refinement of the rule set based on the labeled dataset, using unidentified sub-regions. For example... Figure 8-11 As shown in the example, the two-dimensional XY axis (where K =2), similar to Figure 2 The diagram depicts a sub-region designated by a superrectangular shape representing an approximately continuous curve (702), which represents a basic natural / physical / engineered system.
[0095] exist Figure 8 In the input space, the initially received data points labeled L0 are represented by asterisks (802), (804), (806), and (808). Therefore, the confirmed sub-region is the column sub-region on the left, and the entire remaining portion of the input space is the unconfirmed sub-region of L0 or has a "vertical shading" (810).
[0096] exist Figure 9 In the input space, the L1 data point (902) is received at the upper left corner. To combine this data point (902), as follows... Figure 9 The L1 subregion (910) is created as described in the diagram. Its unconfirmed subregion (910) is reduced in the Y dimension so that it does not overlap with the hyperrectangle corresponding to the confirmed L0 subregion (810), which is defined by the L0 data points directly below it in the Y dimension. Furthermore, the unconfirmed L0 subregion (810) is reduced in the Y dimension so that it does not overlap with the confirmed L1 subregion (910).
[0097] exist Figure 10 In the L0 sub-region, L1 data points (1002) are received, for example, after receiving additional enclosed data points. It is relatively close to... Figure 10 The confirmed subregion of L0 (810) depicted in the diagram / another L0 data point (808). Therefore, the conflict resolution can decide to ignore this data point instead of taking advantage of the fact that the mapping is approximate to create a new subregion. Thus, the subregion containing L0 (810) remains unchanged.
[0098] exist Figure 11 In the middle, receiving the leftmost L1 data point (1102) makes it possible to... Figure 11 The new unconfirmed L1 subregion (1110) shown creates a new subregion pair. The unconfirmed L0 subregion (810) is reduced in the X dimension so that it does not intersect with this new unconfirmed L1 subregion (1110).
[0099] Preprocessing is used to revise attribute values. In one embodiment, the set of input records is preprocessed to improve... Figure 6A and 6BThe process described herein involves automatic rule generation. For example, attributes can indicate the classification or enumeration of temperature ranges, yet make their values inconsistent with semantic proximity. For instance, if the range values—neutral, warm, cool, hot, cold, very hot, and very cold—are assigned sequentially according to the order of these names, it is clear that the value for cool is not between cold and neutral, and similar for several of the other values. Preprocessing can effectively re-assign these numbers to reflect the semantic order, such that each named range has a value that is exactly higher than the temperature range preceding it and lower than the temperature range following it. That is, the value for cool is between cold and neutral, and similar for several of the other values.
[0100] As another form of preprocessing, values can be varied to amplify their dynamic range or adjusted for different sensitivity levels across a range of values. For example, if a temperature sensor has an error limit of 0.5 degrees for temperatures between 32 and 100°F, but a 1.0-degree error limit for temperatures outside that range, then the temperature values between 32 and 100°F can be doubled, and the temperature values outside that range can be added to the resulting minimum and maximum values, that is, 64 to 200. Common error limits can then be used for temperature attributes, regardless of the value. Then, when deploying the resulting rule system, the same preprocessing is applied to ensure consistent results, for example, providing the rule system with the actual temperature of 100°F as the value 200. As another example, when the value of an attribute in the input record is missing or unknown, preprocessing can estimate the value of that attribute.
[0101] Unlabeled datasets can be treated as labeled datasets through preprocessing to compute labels based on one or more attributes selected from the input records, and then one or more of these selected attributes may be excluded from the consideration of rule generation. For example, if one of the attributes is the rate of temperature change in a chemical process, the set of ranges of these rates of change can be identified as labels. Rule generation can then produce rules that substantially capture the association between other input attributes and that rate of temperature change. Other forms of preprocessing transformations can be applied to this set of input records to improve the behavior of the overall system.
[0102] Preprocessing is used to provide time-series records. In one embodiment, the input records have some kind of time indication, such that a subset of the input records corresponding to a sequence of time values is preprocessed into time-series input records. Each such time-series input record indicates a variety of time-related values, wherein the time and values are determined in part based on the original input records. For example, suppose the input records indicate the temperature and level of a furnace at 5-second intervals. Preprocessing can produce time-series records over 60-second intervals, containing 122 tuple-based values corresponding to the temperature and furnace temperature at 5-second intervals for each record. The generated rules may then have antecedents triggered by the time series of states, rather than just combinations of input values at specific times.
[0103] The start time for a time series record can be chosen by changing some input attributes. Continuing with the example above, a new time series record can begin when the furnace's gas level changes significantly. The time series record then captures the furnace's behavior over time based on this change in input.
[0104] Generates based on pre-clustered input records. In one embodiment, the input records are preprocessed to organize them into "clusters," but labels are excluded as part of cluster determination. That is, their attribute values constitute... K Input records of data points that are close to each other in 3D space, independent of their labels, are placed in the same category mentioned in this article. Cluster In this context, clusters can be based on subsets of input attributes. Minimum bounded rectangle (MBR) algorithms or similar algorithms can be used to determine the boundaries around each cluster. As mentioned in this paper, organize This means that preprocessing effectively generates a new set of input records, where all records in a cluster are processed together, separate from other records in other clusters.
[0105] After rearranging the set of input records into a cluster, Figure 6A and Figure 6B The above method is applied to the resulting set of input records from clusters, processing all records in the clusters according to the earlier method, generating sub-regions for each cluster. The processing then checks for overlaps and potential merging between sub-regions across clusters. This final step can be performed after processing for each cluster is complete, or as processing for each cluster occurs, such as after input processing for each cluster is complete.
[0106] In one embodiment, clusters are predefined by sub-regions of the input space. For example, the input space can be divided based on the boundary geometry of each such region. NThere are several regions. Then, if the data points of an input record are within the bounded geometry of that region, the input record is added to the cluster. These regions are then used in traditional bucket sorting methods and, in the sense discussed herein, are essentially buckets. Each such region is called a bucket region. For example, in one embodiment using range-structured rules, a bucket region can be defined by an axis-aligned hyperrectangle. In fact, a bucket is formed by its bounded geometry within a certain range. K A cluster defined by a boundary in dimensional space, independent of the input data points, is the opposite of a cluster defined by the proximity of the input data points.
[0107] Generated based on partially hierarchical input records. In one embodiment, during execution... Figure 6A and 6B Prior to the process described herein, the set of input records is partially hierarchical. In one embodiment, the set of input records is partially hierarchical by distributing each input record into clusters in a cluster sequence, wherein the cluster sequence is ordered according to a certain sorting relation. For example, if the buckets correspond to... K A region in 3D space, as described above as a "bucket region," if the bucket... Bj If it is the next higher bucket area, then the bucket... Bj It's a bucket Bi The successor to this is that if a bucket region contains higher data points, then the bucket region is considered higher, and the data points are considered to be lexicographically ordered according to their attribute values. For example, buckets that partially correspond to temperatures in the range corresponding to cold are ranked so that they appear before buckets that partially correspond to cool temperatures, which are before warm, which are before hot.
[0108] Then, generation continues by processing one bucket at a time, as input records are read and transformed into data points in that input space, using... Figure 6A and Figure 6B A similar process described in the document generates sub-regions within the current bucket region. If the input records in the current bucket have already been processed, processing continues to the next bucket.
[0109] In one embodiment, preprocessed input from the cluster and cluster sorting based on proximity are used. Activity sub-region In this paper, a subregion refers to a point in the process where, during processing, a subsequently generated subregion might be close enough to be merged with it by a proximity metric. Alternatively, a subregion is considered inactive once the input processing is working on a cluster such that no input record data point in that cluster or a subsequent cluster is close enough by a proximity metric to be merged with it. Informally, a subregion in a cluster located only on one far side of the input space may not be merged with a subregion generated on the other side of the input space.
[0110] Therefore, once a sub-region is determined to be inactive, it is not considered for further merging, thus reducing the overhead of checking engulfment / merging. That is, the operation of engulfing / merging into a sub-region of a bucket checks if any sub-region within that sub-region is no longer active, i.e., inactive, and if so, attempts to engulf / merge each such sub-region into the next containing bucket. For example, if any input record after the current input record is too far from bucket region B to contribute to a sub-region that could potentially be merged with a sub-region in B, then B is then determined to be inactive at that level.
[0111] In one embodiment, multiple active clusters exist based on proximity metrics. Subregions deemed inactive in their current cluster are "promoted" to the next cluster. In one embodiment, a hierarchy of clusters exists such that an inactive cluster at one level can be merged into a cluster at the next level, causing the creation or merging of subregions at that level with subregions in that next-level cluster.
[0112] Automatic bucket generation method. Figure 12A and 12B This is a flowchart illustrating an embodiment of a process for automatically approximating and refining a rule set by bucketing based on a labeled dataset. In one embodiment, Figure 12A and 12B The process is Figure 1 The system processing. In one embodiment, Figure 12A and 12B The process is Figure 6A and 6B An enhanced version of the process.
[0113] In step (1202), the initial rule set and associated attribute information are input, including the number of attributes ( K The initial rule set is a boundary-structured rule set, or it can be transformed / converted into labeled sub-regions.
[0114] In step (1204), the set of activity buckets is initialized according to the input rule set, including zero. Current bucket areaAnd a set of its associated sub-regions. Rules in the input rule set are mapped onto these buckets. For example, there may be a super bucket corresponding to the entire input space and a corresponding set of sub-regions containing the sub-regions corresponding to these rules. In step (1206), the next input record is read. In step (1208), if no more input records are available, then in step (1210), any additional merging is optionally performed to convert the set of sub-regions into rules, and then these rules are output as the resulting output rule set; otherwise, control is transferred to step (1212).
[0115] In step (1212), if the available input record is part of the current bucket region, control transfers to step (1214); otherwise, control transfers to step (1222). In step (1212), determining whether an input record is in the current bucket region includes determining whether the data point specified by the input record is in the current bucket region. It is assumed that there exists a way to determine the bucket region for a given input record. As a simple example, the input space can be divided into hyperrectangles of fixed width, effectively... K A 3D grid, therefore a bucket region for a given input is the bucket region of the grid cells to which the data points associated with that input are mapped.
[0116] In step (1214), if the input record as a data point can be merged into an existing sub-region in the current bucket, that is, within the approximate boundary of the sub-region, and the existing sub-region can be expanded to include the data point without conflicting with other sub-regions, control is transferred to step (1218); otherwise, control is transferred to step (1216).
[0117] In step (1218), if the input record has the same or similar label as the existing sub-region identified in step (1214), control is transferred to step (1220); otherwise, control is transferred to step (1222). In step (1220), the sub-region of the input record is expanded to encompass the input data point.
[0118] In step (1216), the new current bucket region begins with the sub-region corresponding to the input record / new input. In step (1222), any sub-region of the current bucket is encompassed / merged into the active bucket sub-region and / or the union of the active bucket sub-regions.
[0119] exist Figure 12A and 12BIn one embodiment not shown, steps (1214) and (1218) are performed in reverse order, i.e., step (1218) is performed before step (1214). More importantly, all three conditions of steps (1212), (1214), and (1218) are true in order to execute step (1220). In one embodiment, one or more buckets are processed in parallel, rather than in a strictly sequential manner. In one embodiment, each bucket region is a hyperrectangle. Therefore, multiple bucket regions are merged only if the result is a hyperrectangle or a collection of hyperrectangles.
[0120] In one embodiment, with attribute A The width of the hyperrectangle in the corresponding dimension is restricted to an attribute. A The approximate limit is about half of the limit. For example, if the approximate limit for temperature is 10 degrees, then there are buckets for 0, 5, 10, 15, 20, etc. Therefore, each time period spans these values, that is, it has values... V Bucket processing A Scope ,in It is an attribute A The approximate limit is one-quarter. Therefore, a temperature of 7.5 can be used in 10 barrels, while a temperature of 17.499 can be used in 15 barrels. Therefore, the values in adjacent barrels are at most separated by an approximate limit, and... A Values in non-adjacent buckets are too far apart to be merged.
[0121] This is generated based on pre-classified input records. In one embodiment, as a special case of partial pre-classification, the input records can be fully classified based on the classification of each attribute. For example, if X1 is less than X2, the input record with the first dimension value X1 appears before the input record with the first dimension value X2. Similarly, this can be extended to the second dimension, and so on. This can be considered as imposing a lexicographical order on the input processing. Using this classification, the last attribute in the ranking is effectively the "least important part," similar to a normal positional notation, where here, "least important" is specified only by the ranking; the second to last is the second least important part, and so on.
[0122] Therefore, the initial bucket corresponds to each combination of attribute values that excludes the last attribute. Thus, logically it is... K A "strip" is a normal measurement of the first dimension up to the second. K Each of the -1 dimensions has a single value.
[0123] Based on this hierarchy, the process effectively iterates over all input records with the same value in the first to the second-to-last attribute, proceeding in ascending order of values in the last attribute before reaching a new value for the second-to-last attribute (which becomes the next bucket). This final sequence of attributes is referred to in this paper as... strip Because it is basically embedded in K A 1-D subregion in 3D space, i.e., a strip. During strip processing, it generates one or more of what are referred to in this paper as strips. Sub-strip bring ,These Sub-strip Each sub-strip in the table is a subset of data points from the strip that have the same or similar labels and are within approximately the same threshold. These sub-strips correspond to sub-regions of the current bucket in the previous method.
[0124] For the generation of pre-graded data, it is sufficient to check the following using each subsequent input record: 1. Does this record initiate a new stripe? 2. Does the record have a label that is significantly different from previous records; or 3. Whether the record's last attribute value is outside the approximate threshold relative to previous records.
[0125] If any of these three conditions is true, the current sub-strip is terminated and added to the sub-strip set. If it is starting a new stripe, the sub-strip set is processed and emptied by either merging these sub-stripes into an existing 2-D sub-region, or creating a new corresponding 2-D sub-region and removing these sub-stripes from the sub-strip set. The new sub-strip is then started with this new input record. Alternatively, sub-stripes of a stripe can be processed into 2-D sub-regions at generation time.
[0126] Here, a 2-D subregion refers to a subregion where the first K-2 attributes are equivalent to a single value or range. If this 2-D subregion is generated by merging two or more 1-D subregions, the penultimate and last attribute values are typically non-singular ranges, although they may be singular ranges. A subregion is considered to be of the dimension of the subregion if the last input record has an attribute value that is within the approximate bound of the maximum value of that attribute in the subregion. activities Otherwise, it is considered in this article. inactive In other words, proximity measures operate along the dimensions (from the last dimension to the first dimension).
[0127] A sub-strip is merged into an existing 2D sub-strip if it meets the following conditions: the sub-strips have approximately the same label; for the first to the penultimate attribute / dimension, the sub-strips have approximately the same range; the penultimate dimension / attribute value is within the approximate bounds of the sub-strip; and the last attribute value of the sub-strip and the last attribute value of the sub-region can be adjusted to be within the approximate bounds of that attribute. Merging 2D sub-regions into 3D sub-regions continues similarly, as is done for merging 3D sub-regions into 4D sub-regions, and so on, until the... K Dimensions and including the first K Dimension. In these cases, merging requires that the sub-regions be approximately identical in the previous dimension and within an approximate threshold in the current dimension.
[0128] An automatic method for generating pre-graded classifications. Figure 13A and 13B This is a flowchart illustrating an embodiment of an automatic approximate refinement process for pre-grading a rule set based on a labeled dataset. In one embodiment, Figure 13A and 13B The process is Figure 1 The system processing. In one embodiment, Figure 13A and 13B The process is Figure 6A and 6B An enhanced version of the process.
[0129] In step (1302), the initial rule set and associated attribute information are input, including the number of attributes ( K The initial rule set is a boundary-structured rule set, or it can be transformed / converted into labeled sub-regions.
[0130] In step (1304), a set of initial dimensional active sub-regions is defined based on the input rule set. Rules in the input rule set are mapped onto these buckets, and as described above, these buckets correspond to each combination of attribute values that excludes the final attribute. Therefore, it is logically... K A "strip" is a normal measurement of the first dimension up to the second. K Each of the -1 dimensions has a single value.
[0131] In step (1306), the next input record is read. In step (1308), if no more input records are available, then in step (1310), the set of active sub-regions is converted into rules, and these rules are then output as the resulting set of output rules; otherwise, control is transferred to step (1312).
[0132] In step (1312), if the available input record is part of the same stripe region, control is transferred to step (1314); otherwise, control is transferred to step (1322). In step (1312), determining whether an input record is part of the same stripe includes determining whether the data point specified by the input record is in the same stripe.
[0133] In step (1314), if the input record is outside the approximate threshold relative to the previous record in terms of its last attribute value, control is transferred to step (1318); otherwise, control is transferred to step (1316).
[0134] In step (1318), if the input record has the same or similar label as the existing sub-region identified in step (1314), control is transferred to step (1320); otherwise, control is transferred to step (1322). In step (1320), the last attribute range of the current sub-strip is expanded to include the new last attribute value, and control returns to step (1306).
[0135] In step (1316), the new sub-strip begins from the sub-region corresponding to the input record / new input. As described above, the sub-strip set is processed and cleared by either merging these sub-strips into an existing 2-D sub-region, or creating a new corresponding 2-D sub-region and removing these sub-strips from the sub-strip set. Then, a new sub-strip begins with this new input record. Alternatively, the sub-strips of a strip can be processed into 2-D sub-regions during generation.
[0136] In step (1322), any sub-regions of the current sub-strip are encompassed / merged into the set of active sub-strips and / or active sub-regions. As described above, a sub-strip is merged into an existing 2-D sub-region if the sub-strips have approximately the same label, approximately the same range from the first to the penultimate attribute / dimension, the penultimate dimension / attribute value is within the approximate bounds of the sub-region, and the last attribute value of the sub-strip and the last attribute value of the sub-region can be adjusted to be within the approximate bounds of that attribute.
[0137] review Figure 13A and 13B The pre-grading process is similar to partial or bucket grading methods, but its special features are: 1. Each initial bucket corresponds to a "strip", and the input records are sorted within each strip; 2. The initial buckets or stripes are sorted lexicographically according to their associated attribute values, excluding the last attribute; and 3. The buckets / sub-regions of the i-th dimension are merged into the buckets of the (i+1)-th dimension, until the _i_th ... K There are dimensions, where the input space is K Via.
[0138] The merge operation for active 2-D subregions checks if any of these subregions is no longer active, that is, inactive, and if so, promotes each such subregion to the set of active 3-D subregions (if...). K If the number is greater than 2, it can either be merged into an existing 3-D subregion or a new corresponding 3-D subregion can be created. Each subsequent dimension set promoted to the active subregion will also merge any subregions that are no longer active into the subsequent dimension of the active subregion (if it is not the final (i.e., the first) dimension). For the final dimension, each K Dimensional subregions can optionally be merged with other subregions before being transformed into rules and output to the result rule set.
[0139] In one embodiment, the preprocessing of clustering input records into buckets and the lexicographical ordering of the buckets partly mean that during processing, an input record is in the current bucket or the next bucket. This reduces the number of superregions to consider when determining whether an input is included. It also reduces the number of superregions to consider when merging superregions. For example, in the case of two dimensions, regions are built incrementally from the bottom left to the top right. This also means that in the part of the multidimensional space containing unprocessed buckets, many unidentified regions extend to the input space constraints, making them simpler / easier to update as part of the processing.
[0140] Bucketing allows sets of input records that are close together to be processed together, reducing processing and space overhead, especially for region merging.
[0141] Back Figure 6A and Figure 6B The illustration depicts sub-regions generated using a pre-graded 2-D hyperrectangular boundary structured rule system. For each value of X, stripes are generated along the Y dimension. 1-D sub-strips with the same label and matching dimension are merged to form the leftmost L0 sub-region (704), and similarly for the leftmost L1 sub-region (710). The illustrated L0 and L1 sub-regions are the result of the following: an extended sub-region is generated for each input record, each is merged into a sub-strip in the Y dimension, and then each sub-strip is merged into a 2-D sub-region along the X dimension.
[0142] If there exists a basic K The dimensional subregion, through the use of restricted inheritance, is currently generated K Wei or derivedSubregions override base subregions, so derived subregions do not need to be merged with base subregions. Instead, rule evaluation is used to select the most derived rule during rule evaluation.
[0143] In one embodiment, based on applied knowledge and experimentation, other hierarchical heuristics are applied to reduce the overhead of approximate checks and the number of labeled subregions, and thus reduce the number of BSRs. Furthermore, a pass-through preprocessing of the input data, in addition to rearranging the set of input records to optimize the generation process, can provide statistics about the dataset, such as the number of records with a given label, the size of these clusters, and other information.
[0144] Attribute generation is achieved using reordered attributes. In one embodiment, attributes in the input records are reordered before processing to minimize the cost of processing the input records into subregions. For example, if an attribute in the original input records is an indication of whether the system is powered, all records for systems where power is set to false might have the same label. Therefore, processing the power system indication as the last attribute might be more efficient, as there would be at most a single active subregion for the dimension where the last attribute is false. This can be achieved using reordering if it is not the last attribute in the structured manner of the input dataset records.
[0145] As another example, even if the timestamp attribute is not specified as the first attribute, it can still be important to sort the inputs by the time they occurred. By reordering the attributes before grading and having the pre-grading treat the timestamp attribute as the first attribute, grading is as if the input record format... The property is set to the first property. The generation method uses the same reordering, so the generation occurs as if... The attribute is the same as the first attribute in each input record.
[0146] The same approach applies if pre-grading uses different orders based on attributes. For example, if an embodiment considers the first attribute to be the least important during pre-grading, then in cases where the generation method handles this different ordering, the attributes can be reordered to have the desired effect, such as making the selected attribute the first attribute instead of the last attribute.
[0147] In one embodiment, buckets are used for partial hierarchical classification, where bucketing depends on attribute ordering, and attributes can be reordered as described above to improve performance.
[0148] In one embodiment, for both pre-sorting and generation purposes, reordering is achieved by arranging attributes, rather than actually rewriting all input records to a revised order before generation. That is, reordering is applied as part of pre-sorting when reading input records, and as part of rule generation when reading record input, thus achieving the same generation as if all input records were rewritten to a reordered order with attributes.
[0149] Generally, different attribute sorts that minimize the number of intermediate active subregions can be selected based on applied knowledge and experiments.
[0150] Major conflict. The terminology used in this article. Major Conflict The indications allow for the measurement and control of conflict levels on an application-by-application basis, permitting input data logging of inaccuracies, and allowing boundary geometry constraints of the boundary structured rules system to provide better coverage while incorporating levels of inaccuracy.
[0151] If there are credible input data points in the dataset that map to the bounding geometry (BG), but have labels significantly different from those associated with the BG, and are not close enough to another hyperrectangle to be overridden, then the candidate bounding geometry (BG) significantly conflicts with the set of input records. More generally, a significant conflict arises when the extended rules overlap with a critical number of valid input dataset records with inconsistent labels. Depending on the application, this critical number can be one or more. Based on the attribute / feature information provided to the dataset, this critical number can also be attribute-specific.
[0152] In one embodiment, a major conflict is defined as a significant difference between the label of the input record and the label of the extension rule. For example, in an autonomous aircraft, if the extended sub-region has a label indicating a moderate increase in airspeed, while the input data record indicates a label indicating a slightly less than moderate increase in airspeed, then there is no major conflict. However, if the input data record indicates a moderate decrease in airspeed, then it is a major conflict.
[0153] Attribute / feature information can specify a threshold for this difference, which can also distinguish between negative and positive differences. For example, having this difference at a higher airspeed compared to the airspeed specified in the input records is more acceptable than having it at a lower airspeed. Therefore, continuing with the autonomous aircraft example, the generated rules have a stall condition indication that appears at an airspeed slightly higher than the airspeed at which the aircraft actually stalls, so that the rules are matched to the set of input records within the constraints of the bounded representation and the desired number of rules in the output. After all, in this example, the aircraft should not fly close to the stall speed, and should be controlled away from the stall speed even if the speed becomes close to it. This avoidance of singularities is common in many other applications. Furthermore, in many control and classification applications, approximate correctness is sufficient, as corrections normally occur during subsequent time steps.
[0154] In one embodiment, it is conventional for the categorization attribute to be specified as an enumeration of named values. However, the enumeration values can be assigned such that proximity in the enumeration order is meaningful in the application. Proximity semantics can be indicated in the attribute information of the feature. For example, the functioning status of monitoring equipment can be indicated by categories notFunctioning, minimalFunction, baseFunction, and fullFunction, assigned values of 0, 1, 2, and 3 respectively. Thus, in conflict detection, the input record indicating the minimum function can be overwritten with R in the overlapping area, which overlaps with the rule R indicating the baseFunction, similar to a derived rule. In this way, the attribute information can indicate the preferred conflict resolution. Here, when an input record indicates a conflict, it is assumed that fewer functions rather than more functions are better. This allows for the generation, for example, of a preference for false positives over false negatives when diagnosing a problem. That is, it is better to incorrectly indicate a fault than to ignore symptoms and miss the fault.
[0155] In one embodiment, the conflict solution collects nearby input data points of the conflict. Conflict sub-regions The operation then uses the same approximate bounds as described previously. Conflicting subregions with too few data points are then eliminated, while those with a sufficient number of conflicting subregions may cause existing subregions to be reduced to resolve the conflicts.
[0156] Potential conflicts between input records and base subregions are considered as input record subregions that overwrite the base subregion label, as described earlier regarding restricted inheritance. Note that this... Conflict Solutions It occurs during rule generation, not during rule evaluation as is seen in conventional rule conflict solutions.
[0157] Multiple levels of acknowledgment. In one embodiment, instead of simply having acknowledgment and unacknowledgment, the process can maintain multiple levels of acknowledgment, where "unacknowledgment" corresponds to level 0. After all, "unacknowledgment" merely means that it is not fully supported by the input data points. A sub-region supported only by a small number of data points can be considered partially acknowledgmented.
[0158] Subregions with different acknowledgment levels are expanded and adjusted as described above, except that higher-level acknowledgments can override lower-level acknowledgments. This can be viewed as a generalization of acknowledging subregions overriding unacknowledging subregions. For example, a subregion R1 with 100 data points and one label can override a subregion R2 with a different label and only 3 data points. In this case, subregions associated with different acknowledgment levels are preserved so that subsequent received data points can be adjusted accordingly for overriding. To illustrate, continuing the example above, if subsequent 200 data points are mapped to a subregion R2 with a different label, that subregion might rise above R1 in acknowledgment and then override R1.
[0159] Feature / attribute information is used in selection and conflict assessment. The above description assumes that a set of features has been selected and that the set of features corresponds to the values specified in the input records. Various conventional techniques can be used to select the feature values to be collected from the set of input records.
[0160] Conversely, the available set of input records specifies the set of possible features or attributes that can be used. Starting with this set of potential features, as part of the processing, it is actually done to specify feature fields to be ignored in the set of input records. Furthermore, if the value of a particular attribute does not contribute to the difference in the generated extended rule set, it indicates that the attribute does not provide information gain in the rule set and can be removed. For example, if an attribute indicates that its value does not contribute to the classification, the attribute can be skipped as part of subsequent rule generation, making subsequent generation and / or regeneration more efficient.
[0161] Concurrent execution. The above process can be performed in parallel by subdividing the input space into diverse sub-regions, with each sub-region being generated in parallel by a separate processor. Then, a final stage of processing is performed to merge the sub-regions across these sub-regions. If the final stage is executed in parallel, there may be some duplication of the generated sub-regions. However, this result does not change the output of the rule system using the output rule set. Furthermore, as an optimization, duplicate rules can be removed.
[0162] Figure 14This is a flowchart illustrating an embodiment of the process of automatically approximating and refining a rule set based on a labeled dataset. In one embodiment, Figure 14 The process is Figure 1 The system processing.
[0163] In step (1402), annotation points in the multidimensional space are input based on the set of input records. In an optional step (1404), the annotation points are converted into input boundary structure rules, and the determination (1406) and expansion (1410) steps are applied to the input superregion.
[0164] In one embodiment, the hyperregion is a multidimensional polyhedron. In one embodiment, the hyperregion is a hyperrectangle. In one embodiment, the hyperrectangle is an axis-aligned hyperrectangle, wherein each side of the axis-aligned hyperrectangle is parallel to one of the axes of the multidimensional space.
[0165] In step (1406), it is determined whether the annotation point and / or input superregion is encompassed by a member superregion of the set of superregions specified by the BSRS, where the BSRS includes boundary structuring rules corresponding to the superregions in the multidimensional space. If the annotation point / input superregion is not encompassed in step (1408), control then transfers to step (1410); otherwise, control transfers to stop. In step (1410), the BSRS is expanded or the annotation point / input superregion is eliminated.
[0166] In one embodiment, the BSRS is part of a rule system. In one embodiment, the rule system is a range-structured rule system. In one embodiment, the BSRS includes base rules that can be overridden by one or more boundary-structured rules generated from a set of input records. In one embodiment, extending the BSRS includes extending it with confirmed / unconfirmed boundary-structured rules.
[0167] In one embodiment, in Figure 14 In a step not shown, the set of input records is preprocessed. In one embodiment, in Figure 14 In a step not shown, the set of marked points is grouped within buckets. In one embodiment, in Figure 14 In a step not shown, the set of labeled points is grouped within buckets such that subsets of the set of labeled points within at least one threshold range are grouped within the same bucket. In one embodiment, in Figure 14 In the steps not shown, the set of labeled points is pre-classified.
[0168] In one embodiment, preprocessing transforms the input record set into a time series containing one or more records. In one embodiment, preprocessing includes at least one of the following: clustering and sorting. In one embodiment, preprocessing includes organizing the set of input records into clusters, and wherein determining and expanding includes processing a subset of the clustered annotation points together. In one embodiment, a cluster is predefined as a subregion of a multidimensional space.
[0169] In one embodiment, preprocessing includes organizing the input record set into clusters, wherein the annotations within the clusters are sorted. In one embodiment, the clusters are sorted using a lexicographical sort. In one embodiment, preprocessing includes a general sort of the input record set. In one embodiment, unconfirmed subregions are maintained, wherein unconfirmed subregions are extensions of a given subregion that do not conflict with confirmed subregions. In one embodiment, the extension is the maximum extension.
[0170] In one embodiment, in Figure 14 In the steps not shown, the rule set in the input record set is automatically refined, wherein the automatic refinement includes the input rule set and the rule set is transformed into a set of extended superregions specified by the Extended Boundary Structured Rule Set (ERS); for a given input record (IR) in the set of input records: the given IR is transformed into the corresponding superregion; it is determined whether the corresponding superregion is encompassed by the extended member superregions of the set of extended superregions specified by ERS; if the corresponding superregion is not encompassed, the ERS is expanded at least partially by merging the corresponding superregion or eliminating it; and the ERS is output as the refined rule set.
[0171] Although the foregoing embodiments have been described in detail for clarity of understanding, the invention is not limited to the details provided. Many alternative ways of implementing the invention exist. The disclosed embodiments are illustrative and not restrictive.
Claims
1. A system comprising: A memory configured to store a boundary structured rule set (BSRS), wherein the BSRS includes boundary structured rules corresponding to superregions in a multidimensional space; as well as The processor is configured to: Based on the set of input records, input the labeled points in the multidimensional space; Determine whether the marked point is encompassed by a member superregion of the superregion set specified by the BSRS; as well as If the marked point is not included, either expand the BSRS or eliminate the marked point.
2. The system of claim 1, wherein the processor is further configured to convert the marked points into an input superregion corresponding to the input boundary structuring rules, and wherein determination and expansion are applied to the input superregion.
3. The system according to claim 1, wherein the superregion is a multidimensional polyhedron.
4. The system according to claim 3, wherein the superregion is a superrectangle.
5. The system of claim 4, wherein the hyperrectangle is an axis-aligned hyperrectangle, wherein each side of the axis-aligned hyperrectangle is parallel to one of the axes of the multidimensional space.
6. The system of claim 1, wherein the BSRS is part of a rules system.
7. The system according to claim 6, wherein the rule system is a range-structured rule system.
8. The system of claim 1, wherein the BSRS includes basic rules that can be overridden by one or more boundary structured rules generated from the set of input records.
9. The system of claim 1, wherein extending the BSRS includes extending it using confirmed / unconfirmed boundary structured rule pairs.
10. The system of claim 1, wherein the processor is further configured to preprocess the set of input records.
11. The system of claim 10, wherein the processor is further configured to group the set of labeled points within a bucket.
12. The system of claim 10, wherein the processor is further configured to group the set of labeled points within buckets such that subsets of the set of labeled points within at least one threshold range are grouped within the same bucket.
13. The system of claim 10, wherein the processor is further configured to pre-classify the set of labeled points.
14. The system of claim 10, wherein the preprocessing converts the input record set into a time series containing one or more records.
15. The system of claim 10, wherein the preprocessing includes at least one of the following: clustering and sorting.
16. The system of claim 10, wherein the preprocessing includes organizing the set of input records into clusters, and wherein determining and expanding includes processing a subset of the clustered annotation points together.
17. The system of claim 16, wherein the cluster is predefined as a sub-region of the multidimensional space.
18. The system of claim 10, wherein preprocessing includes organizing the set of input records into clusters, wherein the annotation points within the clusters are sorted.
19. The system of claim 18, wherein the cluster is ordered using a lexicographical order.
20. The system of claim 10, wherein preprocessing includes a total sorting of the input record set.
21. The system of claim 1, wherein an unconfirmed subregion is maintained, wherein the unconfirmed subregion is an extension of a given subregion that does not conflict with a confirmed subregion.
22. The system of claim 21, wherein the expansion is a maximum expansion.
23. The system of claim 1, wherein extending the BSRS includes incorporating the annotation points.
24. A computer program product embodied in a non-transitory computer-readable medium and comprising computer instructions for: Based on the set of input records, input the labeled points in the multidimensional space; Determine whether the labeled point is encompassed by a member superregion of a set of superregions specified by the BSRS, wherein the BSRS includes boundary structuring rules corresponding to the superregions in the multidimensional space; and If the marker is not included, expand the BSRS or eliminate the marker.
25. A method comprising: Based on the set of input records, input the labeled points in the multidimensional space; Determine whether the labeled point is encompassed by a member superregion of a set of superregions specified by the BSRS, wherein the BSRS includes boundary structuring rules corresponding to the superregions in the multidimensional space; and If the marked point is not included, either expand the BSRS or eliminate the marked point.
26. The method of claim 25, further comprising automatically refining the rule set based on the input record set, wherein automatic refinement includes: Input a rule set and transform it into a set of extended superregions specified by the Extended Boundary Structured Rule Set (ERS); For a given input record (IR) in the set of input records: Convert the given IR into the corresponding superregion; Determine whether the corresponding superregion is encompassed by an extended member superregion of the extended superregion set specified by the ERS; If the corresponding superregion is not encompassed, the ERS is extended at least in part by incorporating the corresponding superregion or eliminating the corresponding superregion; as well as The ERS output is a refined set of rules.
Citation Information
Patent Citations
Restrictive inheritance-based rule conflict resolution
US20240086724A1