Method and apparatus for identifying duplicates in distributed control systems (DCS)
The method and apparatus in DCS automate the detection of duplicates by generating index tables with weighted attribute relevance, enabling efficient and consistent management of process object lists, thereby reducing manual effort and improving system integrity.
Patent Information
- Application Number
- PCT/IB2025/055345
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-05-23
- Publication Date
- 2026-02-05
AI Technical Summary
The complexity of maintaining homogeneous process object lists in Distributed Control Systems (DCS) is increased due to manual verification and merging of changes during the engineering and commissioning phases, especially when multiple engineers are involved, leading to inefficiencies in identifying and managing duplicates.
A method and apparatus that generate index tables for imported lists, identify critical and non-critical attributes with assigned weights, and perform fuzzy searches to detect potential duplicates based on match percentages, using a proximity value determined by these weights, thereby automating the identification process.
This approach significantly reduces the manual effort required for duplicate detection, enhances efficiency by providing faster search and retrieval, and ensures data consistency, while maintaining the integrity of the DCS configuration.
Smart Images

Figure IB2025055345_05022026_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR IDENTIFYING DUPLICATES IN DISTRIBUTED CONTROL SYSTEMS (DCS)TECHNICAL FIELD
[0001] The present disclosure relates to Distributed Control Systems (DCS). Particularly, the present disclosure relates to a method and an apparatus for identifying duplicates in a DCS.BACKGROUND
[0002] During engineering and commissioning phase of Distributed Control Systems (DCS), large configuration data typically in excel / spreadsheet format are fed to the DCS by importing the process objects list in the plant. Process objects here are commonly referred to signals, tags, controllers, measurement devices, logs, trends, and operation elements such as graphics that serve in plant engineering as input and output. Configurations of these process objects are needed to make it possible to engineer and control a plant. Process object lists are prepared using standard templates such as NE131 which is NAMUR standard device, User Association of Automation Technology in Process Industries (NAMUR) and Process Automation Device Information Model (PADIM) etc. Importing process object list follows loading the process objects in the plant and mapping the columns to the parameters of the process objects and validating its type and properties in acceptable format. Process object lists are continuously updated during the life cycle of the plant and imported regularly to maintain the plant. The engineer needs to manually verify and merge the changes performed during the maintenance phase of the plant and subsequently run all tests using the merged process list. This adds to the complexity of maintaining the homogenous process object lists when multiple engineers are involved in plant life cycle. Hence there is a need to overcome the problems discussed above.
[0003] The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the invention and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art.SUMMARY
[0004] Disclosed herein is a method of identifying duplicates in a Distributed Control Systems (DCS). The method comprises generating, by a processor, one or more index tables for a listimported from a source. The list is related to one or more industrial plants associated with a DCS. Further, the method comprises identifying within the list: one or more critical attributes and corresponding first relevancy, and one or more non-critical attributes and corresponding second relevancy. The corresponding first relevancy and the corresponding second relevancy is identified based on a first weight value assigned to the one or more critical attributes or a second weight value assigned to the one or more non-critical attributes. Finally, the method comprises identifying one or more potential duplicates in the list by performing a search operation and a retrieval operation on the one or more index tables and one or more pre-stored index tables. The one or more potential duplicates are identified based on a proximity value and when at least one of, match percentage of the one or more critical attributes is greater than first match percentage value, and match percentage of the one or more non-critical attributes is greater than second match percentage value . The first weight value and the second weight value are used to determine the proximity value between the one or more index tables of the list and the one or more pre-stored index tables to identify the one or more potential duplicates. The one or more potential duplicates are provided to a user.
[0005] Further, disclosed herein is an apparatus for identifying duplicates in the DCS. The apparatus comprises a processor and a memory. The memory is communicatively coupled to the processor and stores processor-executable instructions, which on execution, cause the processor to generate one or more index tables for a list imported from a source. The list is related to one or more industrial plants associated with a DCS. Further, the processor identifies within the list: one or more critical attributes and corresponding first relevancy, and one or more non-critical attributes and corresponding second relevancy. The corresponding first relevancy and the corresponding second relevancy is identified based on a first weight value assigned to the one or more critical attributes or a second weight value assigned to the one or more non-critical attributes. Finally, the processor identifies one or more potential duplicates in the list by performing a search operation and a retrieval operation on the one or more index tables and one or more pre-stored index tables. The one or more potential duplicates are identified based on a proximity value and when at least one of, match percentage of the one or more critical attributes is greater than first match percentage value, and match percentage of the one or more non-critical attributes is greater than second match percentage value. The first weight value and the second weight value are used to determine the proximity value between the one or more index tables of the list and the one or more pre-stored index tables to identifythe one or more potential duplicates. The one or more potential duplicates are provided to a user.
[0006] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWINGS
[0007] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The same numbers are used throughout the figures to reference like features and components. Some embodiments of the system and / or methods in accordance with embodiments of the present subject matter are now described, by way of example only, and regarding the accompanying figures, in which:
[0008] FIG. 1A shows an exemplary architecture for an apparatus for identifying duplicates in a Distributed Control Systems (DCS), in accordance with some embodiments of the present disclosure;
[0009] FIG. IB shows an exemplary process flow illustrating initial import of a list in Distributed Control Systems (DCS), in accordance with some embodiments of the present disclosure;
[0010] FIG. 1C shows an exemplary process flow illustrating update operation performed in Distributed Control Systems (DCS), in accordance with some embodiments of the present disclosure;
[0011] FIG. ID shows an exemplary process of generating index table from one or more documents, in accordance with some embodiments of the present disclosure;
[0012] FIG. 2 shows a detailed block diagram of an apparatus, in accordance with some embodiments of the present disclosure;
[0013] FIG. 3 shows a flowchart illustrating a method of identifying duplicates in a Distributed Control Systems (DCS), in accordance with some embodiments of the present disclosure; and
[0014] FIG. 4 illustrates a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
[0015] It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether such computer or processor is explicitly shown.DETAILED DESCRIPTION
[0016] In the present document, the word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment or implementation of the present subject matter described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0017] While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the specific forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure.
[0018] The terms “comprises”, “comprising”, “includes”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device, or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises. . . a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or method.
[0019] In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by wayof illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense.
[0020] FIG. 1A shows an exemplary architecture for an apparatus for identifying duplicates in a Distributed Control Systems (DCS), in accordance with some embodiments of the present disclosure.
[0021] Exemplary architecture 100 comprises an apparatus 101 and a Distributed Control Systems (DCS) 103. The DCS 103 may comprise a repository 105 which may store one or more index tables which may be generated from a list imported from a source. As an example, the source may be an operator associated with the DCS 103. The list may be related to one or more industrial plants associated with the DCS 103. As an example, the one or more industrial plants may be, without limitation, manufacturing plants, chemical plants, refinery plants, power plants and the like. In an embodiment, the DCS 103 imports list such as process objects lists and overwrite the existing configuration in the one or more industrial plants. The imported list may be verified and the changes in the list may be merged with the existing list in the repository 105. In an embodiment, the apparatus 101 and the DCS 103 may be connected using the communication network (not shown in figure). As an example, the communication network may be a wired communication network, a wireless communication network or a combination of both, which enables the connection of the apparatus 101 and the DCS 103 for communication. In some embodiments, the existing DCS 103 may be configured to perform the functionality of the apparatus 101 as discussed in the present disclosure.
[0022] In an embodiment, the apparatus 101 may be configured to generate the one or more index tables for the list imported from the source. In an embodiment, the apparatus 101 may generate an inverted table index for the imported list which may be an index structure that maintains two hash indexed tables. One of the hash indexed table may be generated for documents and another hash indexed table may be generated for terms within those documents. The document table may include, without limitation, a set of document records, each including the document Identifier (ID) and a list of terms (or pointers to terms) that are present in the document, arranged according to a predefined relevance measure. The term table includes a set of term records, each including the term ID and list of document IDs in which the term occurs.The process of indexing may include preprocessing of text to homogenize the data and remove noise and tokenization of this text to identify individual terms. Further, the data is parsed and stored in the inverted table index based on the indexing scheme. Typical schemes for indexing may include, without limitation, stop word removal, lower case translation, stemming and the like.
[0023] As shown in FIG. IB, step 121, the list is imported from the source and at step 123, the one or more index tables for the list imported is generated. The generated one or more index tables are stored in the repository 105 (step 123). This operation may be performed when the list was performed for the first instance. In some embodiments, this operation may also be performed while updating the one or more index tables in the repository 105. In an embodiment, the one or more index tables may include a first table related to one or more documents including the list and second table related to a plurality of terms in the one or more documents. As shown in FIG. ID index table (also referred as invented index table) is generated from three documents i .e . , document 1 , document 2 and document 3. The index table may store terms in the one or more documents and corresponding document ID storing those terms.
[0024] In an embodiment, upon generating the one or more index tables, the apparatus 101 may be configured to identify within the list: one or more critical attributes and corresponding first relevancy, one or more non-critical attributes and corresponding second relevancy. In an embodiment, the one or more critical attributes and the one or more non-critical attributes may be provided by the user / operator. As an example, the one or more critical attributes may include Input / Output (I / O) signals, direction (input or output), setpoints, instrument type, temperature ranges, pressure thresholds, units and calibration ranges. Similarly, the one or more non-critical attributes in steel manufacturing industry may include alarm triggers, Piping and Instrumentation Diagram (P&ID) references and cabinet / slot information. The apparatus 101 may identify the corresponding first relevancy and the corresponding second relevancy based on a first weight value assigned to the one or more critical attributes and a second weight value assigned to the one or more non-critical attributes. In an embodiment, as the imported list may include different attributes, several fields, corresponding to the attributes of the process element. Some of these attributes are more important to the search, and thus are deemed to be critical attributes. The other attributes for the process element are deemed to be non-critical elements. The critical attributes, are assigned a higher weight than non-critical attributes. Theseweights are then used by the apparatus 101 to compute the relevance of search results when updating the repository 105.
[0025] In an embodiment, upon identifying the one or more critical attributes and the corresponding first relevancy, and the one or more non-critical attributes and the corresponding second relevancy, the apparatus 101 may be configured to identify one or more potential duplicates in the list by performing a search operation and a retrieval operation on the one or more index tables and one or more pre-stored index tables (As shown in FIG. 1C). The search operation may be a fuzzy search. The one or more pre-stored index tables may be previously imported tables which are stored in a repository 105 associated with the DCS 103. In an embodiment, each row in the imported list may be parsed according to the indexing scheme to retrieve individual attributes. These attributes are searched against the corresponding fields in the index to find possible match within the one or more pre-stored index tables. An exact match with the field may return 100%, while an approximate match returns the proximity to the field being searched with. As an example, performing the fuzzy search for the text ‘AA10234567’ against the text ‘AAlo234567’ yields a match of 90%, as the two fields differ by only one character out of 10.
[0026] The apparatus 101 may identify the one or more potential duplicates based on a proximity value and when at least one of, match percentage of the one or more critical attributes is greater than first match percentage value, and match percentage of the one or more non- critical attributes is greater than second match percentage value. The first weight value and the second weight value may be used to determine the proximity value between the one or more index tables of the list and the one or more pre-stored index tables to identify the one or more potential duplicates. In an embodiment, the apparatus 101 may determine a divergence between the one or more index tables of the list and the one or more pre-stored index tables in a repository 105. The one or more potential duplicates are provided to a user. Table A below shows exemplary values of the one or more potential duplicates identified by the apparatus 101.Table B
[0027] As shown in Table B above, the object name stored in the repository and the object parameter 2 stored in the repository do not match with the object name in the imported list and the object parameter 2 stored in the imported list. The one or more potential duplicates may be identified based on a proximity value and when at least one of, match percentage of the one or more critical attributes is greater than first match percentage value, and match percentage of the one or more non-critical attributes is greater than second match percentage value. Upon identifying the one or more potential duplicates, the apparatus 101 may be displayed on a display device associated with the DCS 103 (not shown in figure) and an action may be performed on the one or more potential duplicates. In an embodiment, the action may include, at least one of updating an existing element in the repository with a new element related to the imported process object list, creating the new element and creating a new record for the imported process object list. The action may performed based on the user input.
[0028] In an embodiment, the fuzzy search may be performed against the one or more prestored index tables in the repository 105 to check for an exact or approximate match with the one or more index tables generated from the imported list. The result of the search is a number indicating the proximity of the field being searched to the field stored in the index. Weights for the specific attributes i.e., the one or more critical attributes and the one or more non-critical attributes may be considered to compute the proximity of the one or more index tables against the one or more pre-stored index tables. The fuzzy search may be performed for each field in the document. Further, the apparatus 101 may assign a proximity score to each field taking into account and the weight assigned to it. The proximity score for the entire document is calculated by taking the aggregate of the proximity score for each field and normalizing the proximity score. The proximity score thus is a number between 0 and 1, with 1 indicating a perfect match, and 0 indicating a complete dissonance between the one or more index tables and the one or more pre -stored index tables in the repository 105.
[0029] In an embodiment, the proximity score may calculated using a modified Levenshtein Distance Algorithm (LDA). The algorithm is defined recursively to calculate the divergencebetween two strings (entries from the spreadsheet and index field respectively). In an embodiment, the LDA may be modified by including the weight of the attribute / field which is being checked for update. The fuzzy matching score is computed through a customized recursive algorithm that takes into account the specific weights associated with the one or more critical attributes and the one or more non-critical attributes. This helps in ensuring that the final fuzzy matching score is indicative of the proximity of the string as well as the criticality of the attribute being compared. In an embodiment, the fuzzy matching score is based upon the Levenshtein distance between strings. The fuzzy matching can also be based on other algorithms, such as Hamming Distance Algorithm or Bitmap Algorithm.
[0030] For instance, the computation of the fuzzy matching score based on the modified Levenshtein distance metric is defined as given below:The Fuzzy matching score between two strings a and b is given by leva,b(len(a), len(b)) leva,b(i, j) is the distance between the first i characters of the string ‘a’ and the string ‘b’, and is equal to: leva,b(i, j) = max(i, j), if min(i, j) = 0 otherwise leva,b(i, j) = min(leva,b(i-l, j) + 1, leva,b(i, j-1) + 1, leva,b(i-l, j-1) + lai^bj) lai^bj is the indicator function, given by= 0, when ai = bj= W, otherwise where W is the weight assigned to the attribute to be compared against the index, based on whether the attribute is the critical attribute or the non-critical attribute.
[0031] In an embodiment, the imported list may be considered as a potential duplicate if each of the one or more critical attributes matches attributes in the one or more pre-stored index tables by greater than or equal to a first threshold value for example 95% and / or each of the one or more non-critical attributes match the attributes in the one or more pre-stored index tables by greater than or equal to a second threshold value for example 90%. The values for the percentage used to determine potential duplicates may be configurable and can be specified by the user / operator. As an example, potential duplicate is detected if each of the one or more critical attributes has a match of 90% or greater and / or each of the one or more non-criticalatributes have a match of 95% or greater. All potential matches that meet these criteria are ranked by order of relevance and displayed on a display device associated with the DCS 103 (not shown in figure). The user / operator may decide whether to update an existing document / process element or create a new element in the repository 105. If none of the potential duplicates are deemed relevant or existing data need not be updated, a new record is created in the repository 105. Using an index-based search, as opposed to searching individual records in the repository 105 allows for a much faster search and retrieval. The entire search and update can be processed in linear time as opposed to polynomial time in the case of a search against the repository 105 elements. The repository 105 update is more efficient when using a search index, as only relevant columns need to be updated. Relevant columns in this case are columns in a particular element that differ in value from the existing object in the repository 105. The speed and efficiency in searching comes at the cost of additional space needed to store the index. Further, data consistency needs to be ensured between the repository 105 and the one or more index tables. This can be done by updating the index every time the user / operator takes a decision to update or add an element to the repository 105. The index can be updated simultaneously whenever the repository 105 elements are updated.
[0032] FIG. 2 shows a detailed block diagram of an apparatus 101, in accordance with some embodiments of the present disclosure.
[0033] In some implementations, the apparatus 101 may include an I / O interface 201, a processor 203 and a memory 205. In an embodiment, the memory 205 may be communicatively coupled to the processor 203. The processor 203 may be configured to perform one or more functions of the apparatus 101 for identifying duplicates in a Distributed Control Systems (DCS) 103, using the data 207 and the one or more modules 209 of the apparatus 101N. In an embodiment, the memory 205 may store data 207.
[0034] In an embodiment, the data 207 stored in the memory 205 may include, without limitation, pre-stored index tables data 211 and other data 213. In some implementations, the data 207 may be stored within the memory 205 in the form of various data structures. Additionally, the data 207 may be organized using data models, such as relational or hierarchical data models. The other data 213 may include various temporary data and files generated by the one or more modules 209.
[0035] In an embodiment, the pre-stored index tables data 211 may store one or more prestored index tables which were previously imported tables. In some embodiments, the prestored index tables data 211 may be stored in a repository 105 associated with the DCS 103 system. In an embodiment, each time when a list is imported, one or more index tables generated from the imported list may override the pre-stored index tables data 211 with the one or more index tables generated from the imported list. In an embodiment, one or more potential duplicates in the list are identified by performing a search operation and a retrieval operation on the one or more index tables of the imported list and one or more pre-stored index tables stored in the repository 105.
[0036] In an embodiment, the data 207 may be processed by one or more modules 209 of the apparatus 101. In some implementations, the one or more modules 209 may be communicatively coupled to the processor 203 for performing one or more functions of the apparatus 101. In an implementation, the one or more modules 209 may include, without limiting to, a generating module 215, a identifying module 217, and other modules 219.
[0037] As used herein, the term module may refer to an Application Specific Integrated Circuit (ASIC), an electronic circuit, a hardware processor 203 (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and / or other suitable components that provide the described functionality. In an implementation, each of the one or more modules 209 may be configured as stand-alone hardware computing units. In an embodiment, the other modules 219 may be used to perform various miscellaneous functionalities on the apparatus 101. It will be appreciated that such one or more modules 209 may be represented as a single module or a combination of different modules.
[0038] In an embodiment, the generating module 215 may be configured for generating one or more index tables for a list imported from a source. The list is related to one or more industrial plants associated with a DCS 103. In an embodiment, the generating module 215 may generate an inverted table index which may be an index structure that maintains two hash indexed tables. One of the hash indexed table for documents and another hash indexed table for terms within those documents. The document table may include, without limitation, a set of document records, each including the document Identifier (ID) and a list of terms (or pointers to terms) that are present in the document, arranged according to a predefined relevance measure. The term table includes a set of term records, each including the term ID and list of document IDsin which the term occurs. The process of indexing may include preprocessing of text to homogenize the data and remove noise and tokenization of this text to identify individual terms. Next, the data is parsed and stored in the inverted table index based on the indexing scheme. Typical schemes for indexing may include, without limitation, stop word removal, lower case translation, stemming and the like.
[0039] In an embodiment, the identifying module 217 may be configured for identifying within the list generated by the generating module 215: one or more critical attributes and corresponding first relevancy, and one or more non-critical attributes and corresponding second relevancy. The identifying module 217 may identify the corresponding first relevancy and the corresponding second relevancy based on a first weight value assigned to the one or more critical attributes or a second weight value assigned to the one or more non-critical attributes. In an embodiment, as the imported list may include different attributes, several fields, corresponding to the attributes of the process element. Some of these attributes are more important to the search, and thus are deemed to be critical attributes. The other attributes for the process element are deemed to be non-critical elements. The critical attributes, are assigned a higher weight than non-critical attributes. These weights are then used by the identifying module 217 to compute the relevance of search results when updating the repository 105.
[0040] In an embodiment, the identifying module 217 may be further configured for identifying one or more potential duplicates in the list by performing a search operation and a retrieval operation on the one or more index tables and one or more pre-stored index tables. The one or more pre-stored index tables may be previously imported tables which are stored in a repository 105 associated with the DCS 103. The identifying module 217 may identify the one or more potential duplicates based on a proximity value and when at least one of, match percentage of the one or more critical attributes is greater than first match percentage value, and match percentage of the one or more non-critical attributes is greater than second match percentage value. The first weight value and the second weight value may be used to determine the proximity value between the one or more index tables of the list and the one or more prestored index tables to identify the one or more potential duplicates. In an embodiment, the identifying module 217 may determine a divergence between the one or more index tables of the list and the one or more pre-stored index tables in a repository 105. The one or more potential duplicates are provided to a user. In an embodiment, the one or more potential duplicates may be displayed on a display device associated with the DCS 103 and an actionmay be performed on the one or more potential duplicates. In an embodiment, the action may include, at least one of updating an existing element in the repository 105 with a new element related to the imported process object list, creating the new element and creating a new record for the imported process object list. The action may performed based on the user input.
[0041] FIG. 3 shows a flowchart illustrating a method of identifying duplicates in a Distributed Control Systems (DCS) 103, in accordance with some embodiments of the present disclosure.
[0042] As illustrated in FIG. 3, the method 300 may include one or more blocks illustrating a method of identifying duplicates in a Distributed Control Systems (DCS) 103 using an apparatus 101 illustrated in FIG. 2. The method 300 may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform specific functions or implement specific abstract data types.
[0043] The order in which the method 300 is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the scope of the subject matter described herein. Furthermore, the method can be implemented in any suitable hardware, software, firmware, or combination thereof.
[0044] At block 301, the method 300 includes generating, by a processor 203 associated with the apparatus 101N, one or more index tables for a list imported from a source. The list is related to one or more industrial plants associated with a DCS 103. The one or more index tables may include, without limitation, a first table related to one or more documents including the list and second table related to a plurality of terms in the one or more documents.
[0045] At block 303, the method 300 includes identifying, by the processor 203, within the list: one or more critical attributes and corresponding first relevancy, and one or more non- critical attributes and corresponding second relevancy. The corresponding first relevancy and the corresponding second relevancy is identified based on a first weight value assigned to the one or more critical attributes or a second weight value assigned to the one or more non-critical attributes.
[0046] At block 305, the method 300 includes identifying, by the processor 203, one or more potential duplicates in the list by performing a search operation and a retrieval operation on the one or more index tables and one or more pre-stored index tables. The search operation may be a fuzzy search. The one or more pre-stored index tables may be previously imported tables which are stored in a repository 105 associated with the DCS 103. The one or more potential duplicates are identified based on a proximity value and when at least one of, match percentage of the one or more critical attributes is greater than first match percentage value, and match percentage of the one or more non-critical attributes is greater than second match percentage value. The first weight value and the second weight value may be used to determine the proximity value between the one or more index tables of the list and the one or more pre-stored index tables to identify the one or more potential duplicates. In an embodiment, the processor determines a divergence between the one or more index tables of the list and the one or more pre-stored index tables in a repository 105. The one or more potential duplicates are provided to a user. In an embodiment, the processor may further display the one or more potential duplicates on a display device associated with the DCS 103 and perform an action on the one or more potential duplicates.
[0047] FIG. 4 illustrates a block diagram of an exemplary computer system 400 for implementing embodiments consistent with the present disclosure. In an embodiment, the computer system 400 may be the apparatus 101 illustrated in FIG. 1A. The computer system 400 may include a central processing unit (“CPU” or “processor” or “memory controller”) 402. The processor 402 may comprise at least one data processor for executing program components for executing user- or system -generated business processes. A user may include a network manager, an application developer, a programmer, an organization or any system / sub-system being operated parallelly to the computer system 400. The processor 402 may include specialized processing units such as integrated system (bus) controllers, memory controllers / memory management control units, floating point units, graphics processing units, digital signal processing units, etc.
[0048] The processor 402 may be disposed in communication with one or more Input / Output (I / O) devices (411 and 412) via I / O interface 401. The I / O interface 401 may employ communication protocols / methods such as, without limitation, audio, analog, digital, stereo, IEEE®-1394, serial bus, Universal Serial Bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, Digital Visual Interface (DVI), high-definition multimedia interface (HDMI),Radio Frequency (RF) antennas, S-Video, Video Graphics Array (VGA), IEEE® 8O2.n / b / g / n / x, Bluetooth, cellular (e.g., Code-Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System For Mobile Communications (GSM), Long-Term Evolution (LTE) or the like), etc. Using the I / O interface 401, the computer system 400 may communicate with one or more I / O devices 411 and 412.
[0049] In some embodiments, the processor 402 may be disposed in communication with a network 107 via a network interface 403. The network interface 403 may communicate with the network 409. The network interface 403 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), Transmission Control Protocol / Intemet Protocol (TCP / IP), token ring, IEEE® 802. 1 la / b / g / n / x, etc.
[0050] In an implementation, the preferred network 409 may be implemented as one of the several types of networks, such as intranet or Local Area Network (LAN) and such within the organization. The preferred network 409 may either be a dedicated network or a shared network, which represents an association of several types of networks that use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / Intemet Protocol (TCP / IP), Wireless Application Protocol (WAP) etc., to communicate with each other. Further, the network 409 may include a variety of network devices, including routers, bridges, servers, computing devices, storage devices, etc. Using the network interface 403 and the network 409, the computer system 400 may communicate with an apparatus 101.
[0051] In some embodiments, the processor 402 may be disposed in communication with a memory 405 (e.g., RAM 413, ROM 414, etc. as shown in FIG. 4) via a storage interface 404. The storage interface 404 may connect to memory 405 including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as Serial Advanced Technology Attachment (SATA), Integrated Drive Electronics (IDE), IEEE- 1394, Universal Serial Bus (USB), fiber channel, Small Computer Systems Interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory devices, solid-state drives, etc.
[0052] The memory 405 may store a collection of program or database components, including, without limitation, user / application interface 406, an operating system 407, a web browser 408, and the like. In some embodiments, computer system 400 may store user / application data 406, such as the data, variables, records, etc. as described in this invention. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle® or Sybase® or PostgreSQL®.
[0053] The operating system 407 may facilitate resource management and operation of the computer system 400. Examples of operating systems include, without limitation, APPLE® MACINTOSH® OS X®, UNIX®, UNIX-like system distributions (E G., BERKELEY SOFTWARE DISTRIBUTION® (BSD), FREEBSD®, NETBSD®, OPENBSD, etc ), LINUX® DISTRIBUTIONS (E G., RED HAT®, UBUNTU®, KUBUNTU®, etc ), IBM® OS / 2®, MICROSOFT® WINDOWS® (XP®, VISTA® / 7 / 8, 10 etc ), APPLE® IOS®, GOOGLE ™ ANDROID ™, BLACKBERRY® OS, or the like.
[0054] The user interface 406 may facilitate display, execution, interaction, manipulation, or operation of program components through textual or graphical facilities. For example, the user interface 406 may provide computer interaction interface elements on a display system operatively connected to the computer system 400, such as cursors, icons, check boxes, menus, scrollers, windows, widgets, and the like. Further, Graphical User Interfaces (GUIs) may be employed, including, without limitation, APPLE® MACINTOSH® operating systems’ Aqua®, IBM® OS / 2®, MICROSOFT® WINDOWS® (e.g., Aero, Metro, etc.), web interface libraries (e g., ActiveX®, JAVA®, JAVASCRIPT®, AJAX, HTML, ADOBE® FLASH®, etc ), or the like.
[0055] The web browser 408 may be a hypertext viewing application. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), and the like. The web browsers 408 may utilize facilities such as AJAX, DHTML, ADOBE® FLASH®, JAVASCRIPT®, JAVA®, Application Programming Interfaces (APIs), and the like. Further, the computer system 400 may implement a mail server stored program component. The mail server may utilize facilities such as ASP, ACTIVEX®, ANSI® C++ / C#, MICROSOFT®, NET, CGI SCRIPTS, JAVA®, JAVASCRIPT®, PERL®, PHP, PYTHON®, WEBOBJECTS®, etc. The mail server may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® exchange, Post Office Protocol(POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, the computer system 400 may implement a mail client stored program component. The mail client may be a mail viewing application, such as APPLE® MAIL, MICROSOFT® ENTOURAGE®, MICROSOFT® OUTLOOK®, MOZILLA® THUNDERBIRD®, and the like.
[0056] Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present invention. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., non-transitory. Examples include Random Access Memory (RAM), Read-Only Memory (ROM), volatile memory, nonvolatile memory, hard drives, Compact Disc (CD) ROMs, Digital Video Disc (DVDs), flash drives, disks, and any other known physical storage media.
[0057] In light of the technical advancements provided by the disclosed method, the claimed steps, as discussed above, are not routine, conventional, or well-known aspects in the art, as the claimed steps provide the aforesaid solutions to the technical problems existing in the conventional technologies. Further, the claimed steps clearly bring an improvement in the functioning of the system itself, as the claimed steps provide a technical solution to a technical problem.
[0058] The terms "an embodiment", "embodiment", "embodiments", "the embodiment", "the embodiments", "one or more embodiments", "some embodiments", and "one embodiment" mean "one or more (but not all) embodiments of the invention(s)" unless expressly specified otherwise.
[0059] The terms "including", "comprising", “having” and variations thereof mean "including but not limited to", unless expressly specified otherwise.
[0060] The enumerated listing of items does not imply that any or all the items are mutually exclusive, unless expressly specified otherwise. The terms "a", "an" and "the" mean "one or more", unless expressly specified otherwise.
[0061] A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the invention.
[0062] When a single device or article is described herein, it will be clear that more than one device / article (whether they cooperate) may be used in place of a single device / article. Similarly, where more than one device / article is described herein (whether they cooperate), it will be clear that a single device / article may be used in place of the more than one device / article, or a different number of devices / articles may be used instead of the shown number of devices or programs. The functionality and / or features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality / features. Thus, other embodiments of invention need not include the device itself.
[0063] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the embodiments of the present invention are intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.
[0064] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.Referral Numerals:
Claims
WE CLAIM:
1. A method of identifying duplicates in a Distributed Control Systems (DCS), the method comprising: generating, by a processor, one or more index tables for a list imported from a source, wherein the list is related to one or more industrial plants associated with a DCS; identifying, by the processor, within the list: one or more critical attributes and corresponding first relevancy, and one or more non-critical attributes and corresponding second relevancy, wherein the corresponding first relevancy and the corresponding second relevancy is identified based on a first weight value assigned to the one or more critical attributes or a second weight value assigned to the one or more non- critical attributes; identifying, by the processor, one or more potential duplicates in the list by performing a search operation and a retrieval operation on the one or more index tables and one or more pre-stored index tables, wherein the one or more potential duplicates are identified based on a proximity value and when at least one of, match percentage of the one or more critical attributes is greater than first match percentage value, and match percentage of the one or more non-critical attributes is greater than second match percentage value, wherein the first weight value and the second weight value are used to determine the proximity value between the one or more index tables of the list and the one or more pre-stored index tables to identify the one or more potential duplicates, and wherein the one or more potential duplicates are provided to a user.
2. The method as claimed in claim 1, further comprises: displaying, by the processor, the one or more potential duplicates on a display device associated with the DCS; and performing, by the processor, an action on the one or more potential duplicates.
3. The method as claimed in claim 1, wherein the one or more index tables comprises a first table related to one or more documents comprising the list and second table related to a plurality of terms in the one or more documents.
4. The method as claimed in claim 1, wherein the search operation is a fuzzy search.
5. The method as claimed in claim 1, wherein the one or more pre-stored index tables are previously imported tables which are stored in a repository associated with the DCS.
6. The method as claimed in claim 1, wherein identifying the one or more potential duplicates comprises determining a divergence between the one or more index tables of the list and the one or more pre-stored index tables in a repository associated with the DCS.
7. A apparatus for identifying duplicates in a Distributed Control System (DCS), the apparatus comprising: a processor; and a memory, communicatively coupled to the processor, wherein the memory stores processor executable instructions, which, on execution, causes the processor to: generate one or more index tables for a list imported from a source, wherein the list is related to one or more industrial plants associated with a DCS; identify within the list: one or more critical attributes and corresponding first relevancy, and one or more non-critical attributes and corresponding second relevancy, wherein the corresponding first relevancy and the corresponding second relevancy is identified based on a first weight value assigned to the one or more critical attributes or a second weight value assigned to the one or more non-critical attributes; identify one or more potential duplicates in the list by performing a search operation and a retrieval operation on the one or more index tables and one or more pre-stored index tables, wherein the one or more potential duplicates are identified based on a proximity value and when at least one of, match percentage of the one or more critical attributes is greater than first match percentage value, andmatch percentage of the one or more non-critical attributes is greater than second match percentage value, wherein the first weight value and the second weight value are used to determine the proximity value between the one or more index tables of the list and the one or more pre-stored index tables to identify the one or more potential duplicates, and wherein the one or more potential duplicates are provided to a user.
8. The apparatus as claimed in claim 7, wherein the processor is further configured to: display the one or more potential duplicates on a display device associated with the DCS; and perform an action on the one or more potential duplicates.
9. The apparatus as claimed in claim 7, wherein the one or more index tables comprises a first table related to one or more documents comprising the list and second table related to a plurality of terms in the one or more documents.
10. The apparatus as claimed in claim 7, wherein the search operation is a fuzzy search.
11. The apparatus as claimed in claim 7, wherein the one or more pre-stored index tables are previously imported tables which are stored in a repository associated with the DCS.
12. The apparatus as claimed in claim 7, wherein to identify the one or more potential duplicates the processor is configured to determine a divergence between the one or more index tables of the list and the one or more pre-stored index tables in a repository associated with the DCS.
Citation Information
Patent Citations
Supporting web-query expansion efficiently using multi-granularity indexing and query processing
US20020059161A1