Data management system, data management method, and data management program

The data management system addresses the challenge of correcting notation variations in data files by using a synonym database to automatically present and apply candidate words, thereby reducing user workload and improving efficiency.

WO2025115444A1PCT designated stage expired Publication Date: 2025-06-05RESONAC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/037229
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-10-18
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing data management systems face challenges in efficiently correcting notation variations in data files, which requires significant user intervention and time.

Method used

A data management system that utilizes a synonym database to identify unknown data item names, calculates their similarity with representative words, presents candidate words to the user, and replaces the unknown words with selected candidate words, thereby reducing the workload of correcting notation variations.

Benefits of technology

The system automates the process of correcting notation variations by presenting users with candidate words based on similarity calculations, significantly reducing the time and effort required for data file corrections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024037229_05062025_PF_FP_ABST
    Figure JP2024037229_05062025_PF_FP_ABST
Patent Text Reader

Abstract

This data management system: acquires a data file including one or more data records related to a composition and a header indicating one or more data item names of the one or more data records; refers to a synonym database storing a correspondence relationship between a representative word and one or more synonyms of the representative word to identify, from the one or more data item names, a data item name different from any of the representative word and the synonyms as an unknown word; for each of one or more representative words stored in the synonym database, calculates the similarity between the unknown word and the representative word; determines at least one of the one or more representative words as at least one candidate word on the basis of the similarity of each of the one or more representative words; presents the at least one candidate word to a user; and, in response to the user selecting one candidate word from the at least one candidate word, substitutes the unknown word with the selected candidate word.
Need to check novelty before this filing date? Find Prior Art

Description

Data management system, data management method, and data management program

[0001] One aspect of the present disclosure relates to a data management system, a data management method, and a data management program.

[0002] There are known mechanisms for correcting spelling variations within data files. For example, Patent Document 1 (JP-A-2005-109523) describes a data name matching processing device that correctly determines the identity of objects when there are spelling variations in the character strings of the objects. This device includes a string similarity calculation processing unit that calculates the similarity of at least two character strings that make up each name, and a string identity determination unit that, based on the string similarity calculation results, determines the identity of at least two character strings that do not exactly match but have a predetermined similarity or higher based on the usage pattern in the document.

[0003] JP 2010-231253 A

[0004] A mechanism is needed to reduce the work of correcting spelling variations in data files.

[0005] A data management system according to one aspect of the present disclosure includes at least one processor, which acquires a data file including one or more data records related to a composition and a header indicating one or more data item names of the one or more data records, wherein at least one of the one or more data item names is a name related to an ingredient or characteristic of the composition, and refers to a synonym database that stores correspondences between representative words and one or more synonyms of the representative words to identify, from the one or more data item names, a data item name that is different from either the representative word or the synonym as an unknown word, calculates the similarity between the unknown word and each of the one or more representative words stored in the synonym database, determines at least one of the one or more representative words as at least one candidate word based on the respective similarities of the one or more representative words, presents the at least one candidate word to a user, and, in response to the user selecting a candidate word from the at least one candidate word, replaces the unknown word with the selected candidate word.

[0006] In this aspect, one or more candidate words for replacing an unknown word not stored in the synonym database are automatically presented to the user based on the similarity between the unknown word and each representative word. Then, in response to the user's selection of one candidate word, the unknown word is replaced with the selected candidate word. Since the user only needs to select one candidate word to correct the data item name, the work of correcting spelling variations within the data file can be reduced.

[0007] According to one aspect of the present disclosure, the work of correcting spelling variations in data files can be reduced.

[0008] FIG. 1 is a diagram illustrating an example of the functional configuration of a data management system; FIG. 2 is a diagram illustrating an example of a synonym database; FIG. 3 is a flowchart illustrating an example of processing by the data management system; FIG. 4 is a diagram illustrating an example of a data file; FIG. 5 is a flowchart illustrating processing for unknown words; FIG. 6 is a diagram illustrating an example of a screen for using the data management system; and FIG. 7 is a diagram illustrating an example of replacing data item names and adding synonym data.

[0009] Various examples of the present disclosure will be described in detail below with reference to the accompanying drawings. In the description of the drawings, the same or equivalent elements are designated by the same reference numerals, and redundant description will be omitted.

[0010] [System Configuration] The data management system according to the present disclosure is a computer system that supports the task of correcting spelling variations in data item names in a data file. The data file includes one or more data records related to compositions and a header indicating the names of one or more data items in the one or more data records. Each data record is a line of data and includes one or more values ​​corresponding to one or more data items. At least one of the one or more data items relates to an ingredient or property of the composition. Thus, at least one of the one or more data item names is the name of an ingredient or property of the composition. In one example, the data management system supports the task of correcting spelling variations in data item names that represent the names of ingredients or properties of the composition. The data management system performs spelling variations on the data item names in the header, rather than on the individual data records.

[0011] In connection with correcting spelling variations in data item names, this disclosure uses the terms "target word," "synonym," and "unknown word." A target word is a word or phrase used as a reference when correcting spelling variations. A synonym is a word or phrase that has the same or nearly the same meaning as a target word, but whose notation, which is a sequence of one or more characters, is different from that of the target word. Synonym data indicating the correspondence between a target word and one or more synonyms of the target word is prepared in advance for the data management system. Therefore, the target word and synonym are known words or phrases to the data management system. An unknown word is a word or phrase that is different from both the target word and the synonym, and therefore is unknown to the data management system.

[0012] A data management system is composed of one or more computers. When multiple computers are used, these computers are connected via a communication network such as the Internet or an intranet to logically construct a single data management system.

[0013] A computer that constitutes a data management system generally comprises a processor, memory, and a communication interface as hardware devices. The processor is, for example, a CPU, and the memory is composed of a flash memory, a hard disk, etc. Each function of the data management system is realized by the processor executing a program stored in the memory. The computer may further comprise input devices such as a keyboard and a mouse, and output devices such as a monitor and speakers.

[0014] A data management program for causing a computer to function as a data management system includes program code for implementing each functional module of the data management system. This data management program may be provided in a state where it is non-temporarily recorded on a tangible recording medium, such as a CD-ROM, a DVD-ROM, or a semiconductor memory. Alternatively, the data management program may be provided via a communications network as a data signal superimposed on a carrier wave. The provided data management program is recorded in memory, for example.

[0015] 1 is a diagram showing the functional configuration of an example data management system 10. In this example, the data management system 10 is connected to a synonym database 21, a composition database 22, and a user terminal 30 via a communication network such as the Internet or an intranet.

[0016] The synonym database 21 is a device that stores synonym data indicating the correspondence between a representative word and one or more synonyms of the representative word. The composition database 22 is a device that stores composition data based on one or more data records of a data file. Both the synonym database 21 and the composition database 22 may be provided in a computer system different from the data management system 10, or may be components of the data management system 10.

[0017] The user terminal 30 is a computer used by a user of the data management system 10. The user terminal 30 may be any of various computers, such as a personal computer, a workstation, a tablet terminal, a smartphone, or a wearable terminal.

[0018] The data management system 10 includes a processor 101. In one example, the processor 101 functions as an acquisition unit 11, a first replacement unit 12, a similarity calculation unit 13, a candidate word determination unit 14, a second replacement unit 15, a synonym registration unit 16, and a composition registration unit 17. The acquisition unit 11 is a functional module that acquires a data file. The first replacement unit 12 is a functional module that replaces data item names that match synonyms in a synonym database 21 with representative words. When at least one data item name is an unknown word, the similarity calculation unit 13, the candidate word determination unit 14, and the second replacement unit 15 each perform processing related to the unknown word. The similarity calculation unit 13 is a functional module that calculates the similarity between an unknown word and each representative word. In the present disclosure, similarity refers to an index that indicates how similar the notations of two words, i.e., two character strings, are. The more similar the two character strings are, the higher the similarity. The candidate word determination unit 14 is a functional module that determines at least one of one or more representative words as a candidate word based on each similarity. The candidate word determination unit 14 presents at least one candidate word to the user. The second replacement unit 15 is a functional module that replaces an unknown word with the selected candidate word (representative word) in response to the user's selection of a candidate word. As a result, the data item name that was an unknown word is replaced with the selected candidate word (representative word). The synonym registration unit 16 is a functional module that registers new correspondences between representative words and synonyms in the synonym database. The composition registration unit 17 is a functional module that registers composition data based on one or more data records in the data file in which synonyms and unknown words have been replaced in the composition database 22.

[0019] FIG. 2 is a diagram showing an example of the synonym database 21. In this example, each data record of the synonym data includes a pair of a representative word and one synonym. FIG. 2 shows a record group 211 related to the representative word "NC3000," a record group 212 related to the representative word "YX4000," and a record group 213 related to the representative word "Tg." The record group 211 shows the correspondence between the representative word "NC3000" and one synonym "NC-3000." The record group 212 shows the correspondence between the representative word "YX4000" and one synonym "YX-4000." The record group 213 shows the correspondence between the representative word "Tg" and two synonyms, "glass transition temperature" and "glass transition point." In the example of Figure 2, to facilitate comparison of the data item names extracted from the data file with the synonym data, a data record is also provided in which the same word is set in both the representative word column and the synonym column for each representative word.

[0020] Other data structures may be adopted for the synonym data as long as they indicate the correspondence between a target word and one or more synonyms. For example, each data record of the synonym data may include a target word and one or more synonyms. In this case, one data record is prepared for each target word.

[0021] The data structure of the composition database 22 is designed to store each data record in the data file as composition data or to store composition data generated based on the data records. The composition database 22 may be configured with a single data table or may be configured with multiple data tables designed by database normalization.

[0022] [System Operation] An example of processing by the data management system 10 and an example of a data management method according to the present disclosure will be described with reference to Fig. 3. Fig. 3 is a flowchart showing this example as a processing flow S1.

[0023] In step S11, the acquisition unit 11 acquires a data file. The acquisition unit 11 may receive a data file transmitted from the user terminal 30. Alternatively, the acquisition unit 11 may access a predetermined storage device in response to a user operation on the user terminal and read the data file from the storage device.

[0024] A data file can be expressed in a table format such as a spreadsheet, CSV, or TSV. Each data record is called a row, and each data item is represented by a column. FIG. 4 shows an example of a data file. The data file 200 shown in this example includes a header 201 indicating the names of multiple data items and one or more data records 202 related to compositions. The header 201 includes, as data items, a "sample name" that is an identifier for uniquely identifying a composition sample; "NC-3000," "YX 4000," and "KBM583" that indicate the raw materials of the composition; and a "Tg" that indicates the properties of the composition. The sample name may be represented by a name or number that identifies the composition, the name or number of the experiment used to obtain the composition, or a string of characters indicating other information related to the composition. The data item name "sample name" is not subject to replacement processing by the data management system 10. The data management system 10 performs substitution processing on data item names relating to ingredients or properties of compositions, such as "NC-3000," "YX 4000," "KBM583," and "Tg."

[0025] Returning to FIG. 3 , in step S12, the first replacement unit 12 performs a first replacement on the data file. The first replacement unit 12 references the synonym database 21 to identify data item names that match synonyms from among one or more data item names in the header. If one or more data item names that match synonyms are identified, the first replacement unit 12 replaces each of the identified data item names with a representative word corresponding to the synonym. This first replacement can change at least one data item name to a representative word. As described above, in the example of FIG. 2 , there are data records in which the same word or phrase is set as both a representative word and a synonym. Therefore, the first replacement unit 12 can also perform the first replacement on data item names that are already indicated by a representative word before the first replacement is performed. However, these data item names remain unchanged before and after the first replacement.

[0026] In step S13, the first replacement unit 12 identifies unknown words from the header of the data file. The first replacement unit 12 refers to the synonym database 21 and identifies, from one or more data item names in the header, data item names that are different from both the representative word and the synonym, as unknown words. That is, the first replacement unit 12 identifies data item names expressed in notations that are not recorded in the synonym database 21. In one example, the first replacement unit 12 identifies unknown words after replacing data item names that match synonyms with the representative word. That is, step S13 can be executed after step S12.

[0027] As shown in step S14, the first replacement unit 12 may identify one or more unknown words, or may not identify any unknown words. If one or more unknown words are identified (YES in step S14), the process proceeds to step S15. If no unknown words are identified (NO in step S14), the process skips step S15 and proceeds to step S16.

[0028] In step S15, the similarity calculation unit 13, the candidate word determination unit 14, and the second replacement unit 15 cooperate to perform a second replacement on the data file, and the synonym registration unit 16 registers the new correspondence between the target word and the synonym in the synonym database. These processes will be described in detail with reference to Figure 5. Figure 5 is a flowchart showing the process for unknown words.

[0029] In step S151, the similarity calculation unit 13 selects one of one or more unknown words.

[0030] In step S152, the similarity calculation unit 13 calculates the similarity between the selected unknown word and the representative word for each of the one or more representative words stored in the synonym database 21. In one example, the similarity calculation unit 13 may calculate the character string distance between the unknown word and the representative word and calculate the similarity based on this character string distance. As one example, the similarity calculation unit 13 may calculate the Levenshtein distance as the character string distance. The Levenshtein distance is expressed as the minimum number of operations required to transform one character string into another, assuming that one operation is the insertion, deletion, or substitution of one character. The smaller the Levenshtein distance between two character strings, the higher the similarity between the two character strings. As another example, the similarity calculation unit 13 may calculate the Jaro-Winkler distance as the character string distance. The Jaro-Winkler distance is calculated based on the number of matching characters between one character string and the other character string and a determination of whether replacement is necessary. The Jaro-Winkler distance is expressed as a value between 0 and 1, and the larger the value, the higher the similarity between the two strings.

[0031] The similarity calculation unit 13 calculates the similarity based on the calculated character string distance. As described above, the relationship between the magnitude of the character string distance and the level of similarity may vary depending on the type of character string distance. When the Levenshtein distance is used, the similarity calculation unit 13 calculates the similarity using an algorithm designed so that the smaller the Levenshtein distance, the higher the similarity. When the Jaro-Winkler distance is used, the similarity calculation unit 13 calculates the similarity using an algorithm designed so that the larger the Levenshtein distance, the higher the similarity. The algorithm may be realized using various methods such as a calculation formula, a correspondence table, or a conversion table. When using a method such as the Jaro-Winkler distance, in which the larger the distance, the higher the similarity between two character strings, the similarity calculation unit 13 may set the calculated distance as the similarity. Setting the similarity in this manner is also an example of calculating the similarity.

[0032] In one example, regardless of the type of character string distance employed, the similarity calculation unit 13 modifies the character string distance to lower the similarity when the numeric string included in the unknown word and the numeric string included in the representative word are different. The similarity calculation unit 13 then calculates the similarity based on the modified character string distance. Regarding the ingredients and properties of compositions, even if two character strings representing two objects are similar overall, if the numeric strings included in the character strings are different, the two objects may be completely different. Taking into account such tendencies in the names of ingredients and properties, the similarity can be calculated more appropriately by modifying the character string distance to lower the similarity when the numeric strings are different. For example, when the unknown word "DM1000" is compared with the representative word "DM1100," the similarity is high because only the fourth character is different. However, in reality, the ingredient called "DM1000" is likely to have completely different properties or characteristics from the ingredient called "DM1100." Therefore, the similarity calculation unit 13 modifies the character string distance between the unknown word "DM1000" and the representative word "DM1100" so that the similarity between these two words becomes lower.

[0033] When the number string included in the unknown word and the number string included in the representative word are different, the similarity calculation unit 13 may apply a penalty of addition, subtraction, multiplication, or division to the calculated string distance to modify the string distance. When the Levenshtein distance is used, the similarity calculation unit 13 may increase the string distance (Levenshtein distance) by adding a predetermined value to the calculated string distance or by multiplying the string distance by a predetermined value. When the Jaro-Winkler distance is used, the similarity calculation unit 13 may decrease the string distance (Jaro-Winkler distance) by subtracting a predetermined value from the calculated string distance or by dividing the string distance by a predetermined value. The similarity calculation unit 13 calculates the similarity based on the string distance modified by such a penalty.

[0034] In step S153, the candidate word determination unit 14 determines at least one of the one or more representative words as at least one candidate word corresponding to the selected unknown word based on the similarity of each of the one or more representative words. For example, the candidate word determination unit 14 may select some of the one or more representative words as candidate words in descending order of similarity. The number of candidate words may be specified by a fixed value such as 1, 2, 3, 10, etc., or may be specified by a percentage (e.g., 10%, 20%, etc.) of the number of representative words indicated by the synonym data.

[0035] As shown in step S154, the similarity calculation unit 13 and the candidate word determination unit 14 cooperate to calculate the similarity and determine candidate words for all unknown words. If there are unprocessed unknown words (NO in step S154), the process returns to step S151. In the repeated step S151, the similarity calculation unit 13 selects the next unknown word. In the repeated step S152, the similarity calculation unit 13 calculates the similarity between the selected unknown word and each of one or more representative words stored in the synonym database 21. In the repeated step S153, the candidate word determination unit 14 determines at least one candidate word corresponding to the selected unknown word based on the similarity of each of the one or more representative words. If all unknown words have been processed (YES in step S154), the process proceeds to step S155.

[0036] In step S155, the candidate word determination unit 14 transmits candidate word data indicating a combination of each unknown word with one or more candidate words (representative words) to the user terminal 30. This transmission is an example of a process of presenting at least one candidate word corresponding to an unknown word to the user. The user terminal 30 receives and displays the candidate word data, allowing the user to confirm the candidate words for each unknown word.

[0037] In step S156, the second replacement unit 15 receives response data indicating each phrase selected for each unknown word from the user terminal 30. The user, who has been presented with the candidate word data, performs one of two operations for each of the one or more unknown words. One operation is to select one from one or more candidate words (representative words) corresponding to the unknown word. The other operation is to select the unknown word as is without selecting a candidate word (representative word). The user terminal 30 generates response data in response to the completion of the operation for each unknown word and transmits the response data to the data management system 10. The second replacement unit 15 receives the response data.

[0038] In step S157, the second replacement unit 15 processes the unknown words based on the response data. The second replacement unit 15 performs the following process for each of the one or more unknown words. If a candidate word (representative word) corresponding to the unknown word is selected, the second replacement unit 15 replaces the data item name that is the unknown word with the selected candidate word (representative word). If the unknown word is selected as is, the second replacement unit 15 maintains the data item name that is the unknown word as is.

[0039] Steps S156 and S157 are an example of a process of replacing an unknown word with a selected candidate word (representative word) in response to a user selecting one candidate word (representative word) from at least one candidate word (representative word).

[0040] In step S158, the synonym registration unit 16 registers a new correspondence between the representative word and the synonym in the synonym database 21. That is, the synonym registration unit 16 adds new synonym data to the synonym database 21. When the user selects one candidate word (representative word) for an unknown word, the synonym registration unit 16 registers the combination of the candidate word (representative word) and the unknown word as a new correspondence between the representative word and the synonym in the synonym database 21. When the user selects an unknown word as is, that is, when the user does not select any candidate word (representative word) for the unknown word, the synonym registration unit 16 registers a new correspondence in the synonym database 21 in which the unknown word is set for both the new representative word and the new synonym.

[0041] Returning to FIG. 3 , in step S16, the composition registration unit 17 registers composition data based on one or more data records of the data file in which the data item names have been replaced in the composition database 22. The "data file in which the data item names have been replaced" may be a data file in which synonyms have been replaced with representative words, or a data file in which unknown words have been replaced with candidate words (representative words) selected by the user. In either case, the "data file in which the data item names have been replaced" is a data file in which spelling variations in the data item names have been corrected. The composition registration unit 17 may process each value in each data record of the data file as composition data as is. Alternatively, the composition registration unit 17 may generate composition data by performing a predetermined calculation on at least one value in the data record. In either case, the composition registration unit 17 registers composition data based on the data file in the composition database 22.

[0042] An example of the processing flow S1 will be described with reference to Figures 6 and 7. Figure 6 is a diagram showing an example of a screen for using the data management system 10. Figure 7 is a diagram showing an example of replacing data item names and adding synonym data.

[0043] In response to the user terminal 30 accessing the data management system 10 based on a user operation, the data management system 10 provides the user terminal 30 with a screen 300. For example, the data management system 10 provides the screen 300 in the form of a web page displayed on a web browser. The screen 300 includes a setting area 310 that accepts uploading of a data file and selection of a character string distance to be used in calculating the similarity, and a selection area 320 that accepts an operation for a data item name identified as an unknown word.

[0044] The user selects one character string distance and uploads the data file. In response to this operation, the data management system 10 executes process flow S1. The following describes process flow S1, assuming the data file 200 and synonym database 21 shown in FIG. 7.

[0045] In response to the acquisition unit 11 acquiring the data file 200 (step S11), the first replacement unit 12 executes a first replacement on the data file 200 (step S12). The first replacement unit 12 identifies, from the header 201, data item names "NC-3000" and "Tg" that match synonyms in the synonym database 21, and replaces "NC-3000" with "NC3000". Because the identified data item name "Tg" is the same as the representative word "Tg," this data item name remains unchanged before and after the first replacement.

[0046] Since the first replacement unit 12 identifies at least two data item names, "YX 4000" and "KBM583," as unknown words (step S13), a second replacement is performed (step S15). For each unknown word, the similarity calculation unit 13 calculates the similarity between each unknown word and each unknown word, such as "NC3000," "YX4000," and "Tg" (steps S151 and S152). The candidate word determination unit 14 determines a candidate word based on the similarity (step S153). The similarity calculation unit 13 calculates the similarity based on the character string distance selected on the screen 300.

[0047] The candidate word determination unit 14 transmits the candidate word data to the user terminal 30 (step S155), and the user terminal displays the candidate word data in the selection area 320. In the example of FIG. 6 , the selection area 320 displays two unknown data item names, "YX 4000" and "KBM583." Each list box displays one or more candidate words (representative words) corresponding to the unknown words. For each unknown word (data item name), the user either selects one candidate word (representative word) from the list box or selects the unknown word (current data item name) as is. In the example of FIG. 6 , the user selects replacing the data item name "YX 4000" with the candidate word "YX4000" or using the data item name "KBM583" as is. The user performs this operation for each unknown word and finally clicks the registration button. In response to this click, the user terminal 30 generates response data and transmits the response data to the data management system 10.

[0048] The second replacement unit 15 receives the response data (step S156) and processes one or more unknown words (step S157). As shown in FIG. 7, the second replacement unit 15 replaces the data item name "YX 4000" with "YX4000" and maintains the data item name "KBM583." The synonym registration unit 16 registers, as new synonym data, a data record 214 indicating the correspondence between the representative term "YX4000" and the synonym "YX 4000," and a data record 215 in which "KBM583" is set as both the representative term and the synonym, in the synonym database 21 (step S158). The composition registration unit 17 registers, in the composition database 22, composition data based on one or more data records in the data file 200 in which the data item names have been replaced (step S16). As a result, composition data relating to the composition samples Qa, Qb, etc., are stored in the composition database 22.

[0049] [Modifications] The technology according to the present disclosure has been described in detail above based on various examples. However, the present disclosure is not limited to the above examples. The technology according to the present disclosure can be modified in various ways without departing from the spirit of the present disclosure.

[0050] The data management system according to the present disclosure may not include at least one of the first replacement unit, the synonym registration unit, and the composition registration unit. That is, a computer system different from the data management system may execute at least one of the processes of replacing data item names that match synonyms with representative words, registering new synonym data, and registering composition data based on a data file.

[0051] In the above example, the first replacement unit 12 identifies the unknown word after replacing the data item name that matches the synonym with the representative word, i.e., after the first replacement. However, the data management system may also identify the unknown word before the first replacement. In this regard, the data management system may also perform the second replacement before the first replacement.

[0052] In the above example, the data management system is implemented as a server in a client-server system. As another example, the data management system may be implemented in a stand-alone computer. Alternatively, the data management system may be implemented in a user terminal that can access predetermined databases, such as a synonym database and a composition database, via a communication network.

[0053] The processing steps of the method executed by at least one processor are not limited to the above examples. For example, some of the above steps may be omitted, or the steps may be executed in a different order. Furthermore, any two or more of the above steps may be combined, or some of the steps may be modified or deleted. Alternatively, other steps may be executed in addition to the above steps.

[0054] In the present disclosure, when comparing the magnitude of two numerical values, either of the two criteria "greater than or equal to" and "greater than" may be used, or either of the two criteria "less than or equal to" and "less than" may be used.

[0055] In the present disclosure, the expression "at least one processor executes a first process, executes a second process, ... executes an nth process" or an expression corresponding thereto indicates a concept including a case where the entity executing the n processes from the first process to the nth process, i.e., the processor, changes midway through. In other words, this expression indicates a concept including both a case where all n processes are executed by the same processor and a case where the processor changes among the n processes according to an arbitrary policy.

[0056] [Additional Notes] As can be seen from the various examples above, the present disclosure includes the following aspects. (Supplementary Note 1) A data management system comprising at least one processor, wherein the at least one processor: acquires a data file including one or more data records related to a composition and a header indicating one or more data item names of the one or more data records, wherein at least one of the one or more data item names is a name related to an ingredient or a property of the composition; refers to a synonym database that stores correspondences between representative words and one or more synonyms of the representative words, and identifies, from the one or more data item names, a data item name that is different from either the representative word or the synonym as an unknown word; calculates, for each of the one or more representative words stored in the synonym database, a similarity between the unknown word and the representative word; determines at least one of the one or more representative words as at least one candidate word based on the respective similarities of the one or more representative words; presents the at least one candidate word to a user; and, in response to the user selecting a candidate word from the at least one candidate word, replaces the unknown word with the selected candidate word. (Supplementary Note 2) The data management system according to Supplementary Note 1, wherein the at least one processor calculates, for each of the one or more representative words, a character string distance between the unknown word and the representative word, and if a numeric string included in the unknown word is different from a numeric string included in the representative word, corrects the character string distance so that the similarity becomes lower, and calculates the similarity based on the corrected character string distance. (Supplementary Note 3) The data management system according to Supplementary Note 1 or 2, wherein the at least one processor registers a combination of the candidate word and the unknown word selected by the user in the synonym database as a new correspondence between the representative word and the synonym. (Supplementary Note 4) The data management system according to any one of Supplements 1 to 3, wherein the at least one processor refers to the synonym database and replaces a data item name that matches a synonym among the one or more data item names with the representative word corresponding to the synonym, and identifies the unknown word after replacing the data item name that matches the synonym with the representative word.(Supplementary Note 5) The data management system according to any one of Supplements 1 to 4, wherein the at least one processor, in response to the user not selecting any of the at least one candidate words, registers in the synonym database a new correspondence in which the unknown word is set for both a new representative word and a new synonym. (Supplementary Note 6) The data management system according to any one of Supplements 1 to 5, wherein the at least one processor, using the data file in which the unknown word has been replaced with the selected candidate word, registers composition data based on the one or more data records of the data file in a composition database. (Supplementary Note 7) A data management method executed by a data management system having at least one processor, comprising: a step of acquiring a data file including one or more data records related to a composition and a header indicating one or more data item names of the one or more data records, wherein at least one of the one or more data item names is a name related to an ingredient or a property of the composition; a step of identifying, from the one or more data item names, as an unknown word, a data item name that is different from either the representative word or the synonym, by referring to a synonym database that stores correspondences between a representative word and one or more synonyms of the representative word; a step of calculating, for each of the one or more representative words stored in the synonym database, a similarity between the unknown word and the representative word; a step of determining at least one of the one or more representative words as at least one candidate word based on the similarity of each of the one or more representative words; a step of presenting the at least one candidate word to a user; and a step of replacing the unknown word with the selected candidate word in response to the user selecting a candidate word from the at least one candidate word.(Supplementary Note 8) A data management program that causes a computer to execute the following steps: acquiring a data file including one or more data records related to a composition and a header indicating one or more data item names of the one or more data records, wherein at least one of the one or more data item names is a name related to an ingredient or a property of the composition; identifying, from the one or more data item names, as an unknown word, a data item name that is different from either the representative word or the synonym, by referring to a synonym database that stores correspondences between a representative word and one or more synonyms of the representative word; calculating, for each of the one or more representative words stored in the synonym database, the similarity between the unknown word and the representative word; determining at least one of the one or more representative words as at least one candidate word based on the similarity between each of the one or more representative words; presenting the at least one candidate word to a user; and replacing the unknown word with the selected candidate word in response to the user selecting a candidate word from the at least one candidate word.

[0057] According to Supplements 1, 7, and 8, one or more candidate words for replacing an unknown word not stored in the synonym database are automatically presented to the user based on the similarity between the unknown word and each representative word. Then, in response to the user's selection of one candidate word, the unknown word is replaced with the selected candidate word. Since the user only needs to select one candidate word to correct the data item name, the work of correcting spelling variations within the data file can be reduced.

[0058] Regarding ingredients and properties of a composition, even if two character strings representing two objects are similar overall, if the numeric strings included in the character strings are different, the two objects are completely different. According to Supplementary Note 2, taking into consideration such tendencies in names of ingredients and properties, when the numeric strings are different, the character string distance is corrected and the similarity is calculated as low, thereby making it possible to more accurately determine candidate words that may be selected by the user.

[0059] According to Appendix 3, pairs of unknown words and target words related to the replacement process are registered in a synonym database, and the unknown words are converted to known words in the data management system. Therefore, the next time the same phrase is replaced with a target word, there is no need to present candidate words to the user. This mechanism further reduces the work required to correct spelling variations in data files.

[0060] According to Supplementary Note 4, unknown words are identified after processing data item names that can be automatically replaced, so that data item name replacement can be performed efficiently overall.

[0061] According to Appendix 5, the unknown word is registered in the synonym database in a form that can be used in subsequent processing, and the unknown word is converted into a known word in the data management system. Therefore, the next time the same phrase is processed, there is no need to present candidate words to the user. This mechanism further reduces the work of correcting spelling variations in data files.

[0062] According to Appendix 6, by using a data file in which variations in the spelling of data item names have been corrected, composition data based on the data file can be smoothly and efficiently registered in the composition database.

[0063] 10...data management system, 11...acquisition unit, 12...first replacement unit, 13...similarity calculation unit, 14...candidate word determination unit, 15...second replacement unit, 16...synonym registration unit, 17...composition registration unit, 21...synonym database, 22...composition database, 30...user terminal, 200...data file, 201...header, 202...data record, 300...screen.

Claims

1. A data management system comprising at least one processor, the at least one processor acquiring a data file including one or more data records relating to a composition and a header indicating one or more data item names of the one or more data records, wherein at least one of the one or more data item names is a name relating to an ingredient or characteristic of the composition, referring to a synonym database storing a correspondence between a representative word and one or more synonyms of the representative word, identifying a data item name which is different from both the representative word and the synonym as an unknown word from the one or more data item names, calculating a similarity between the unknown word and each of the one or more representative words stored in the synonym database, determining at least one of the one or more representative words as at least one candidate word based on the respective similarities of the one or more representative words, presenting the at least one candidate word to a user, and in response to the user selecting a candidate word from the at least one candidate word, replacing the unknown word with the selected candidate word.

2. The data management system of claim 1, wherein the at least one processor calculates a character string distance between the unknown word and the representative word for each of the one or more representative words, and if the numeric string contained in the unknown word and the numeric string contained in the representative word are different, modifies the character string distance so that the similarity is lowered, and calculates the similarity based on the modified character string distance.

3. The data management system according to claim 1 or 2, wherein the at least one processor registers the combination of the candidate word and the unknown word selected by the user in the synonym database as a new correspondence between the representative word and the synonym.

4. A data management system as described in claim 1 or 2, wherein the at least one processor refers to the synonym database, replaces data item names among the one or more data item names that match the synonym with the representative word corresponding to the synonym, and identifies the unknown word after replacing the data item names that match the synonym with the representative word.

5. A data management system as described in claim 1 or 2, wherein the at least one processor, in response to the user not selecting any of the at least one candidate word, registers in the synonym database a new correspondence relationship in which the unknown word is set to both a new representative word and a new synonym.

6. The data management system according to claim 1 or 2, wherein the at least one processor uses the data file in which the unknown words have been replaced with the selected candidate words to register composition data based on the one or more data records of the data file into a composition database.

7. A data management method executed by a data management system having at least one processor, comprising: a step of acquiring a data file including one or more data records related to a composition and a header indicating one or more data item names of the one or more data records, wherein at least one of the one or more data item names is a name related to an ingredient or characteristic of the composition; a step of identifying, from the one or more data item names, a data item name which is different from both the representative word and the synonym, as an unknown word, by referring to a synonym database which stores a correspondence between a representative word and one or more synonyms of the representative word; a step of calculating a similarity between the unknown word and each of the one or more representative words stored in the synonym database; a step of determining at least one of the one or more representative words as at least one candidate word based on the respective similarities of the one or more representative words; a step of presenting the at least one candidate word to a user; and a step of replacing the unknown word with the selected candidate word in response to the user selecting a candidate word from the at least one candidate word.

8. A data management program that causes a computer to execute the following steps: acquiring a data file including one or more data records related to a composition and a header indicating one or more data item names of the one or more data records, wherein at least one of the one or more data item names is a name related to an ingredient or characteristic of the composition; referring to a synonym database that stores a correspondence between a representative word and one or more synonyms of the representative word, and identifying, from the one or more data item names, a data item name that is different from both the representative word and the synonym as an unknown word; calculating, for each of the one or more representative words stored in the synonym database, a similarity between the unknown word and the representative word; determining at least one of the one or more representative words as at least one candidate word based on the respective similarities of the one or more representative words; presenting the at least one candidate word to a user; and, in response to the user selecting a candidate word from the at least one candidate word, replacing the unknown word with the selected candidate word.

Citation Information

Patent Citations

  • Unregistered word acquiring system

    JP1994195371A

  • Method and device for processing document

    JP1997006782A

  • Method and device for correcting japanese character recognition error and recording medium with error correcting program recorded

    JP1999328317A

  • Analysis support method, analysis support server and storage media

    JP2019109676A

  • Standard item name setting device, standard item name setting method, and standard item name setting program

    JP2020004373A