Apparatus and method for processing, analyzing, and visualizing scientific and technological raw data file
By standardizing and unifying multiple names for countries, companies, and universities, the system addresses data distortion, enabling accurate and intuitive data analysis and trend assessment.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- AMUR CO LTD
- Filing Date
- 2025-11-14
- Publication Date
- 2026-05-21
AI Technical Summary
Existing data analysis systems face challenges in accurately processing and analyzing raw scientific and technological data due to the use of multiple names for countries, companies, and universities, leading to data distortion and difficulty in understanding relative trends.
A system that standardizes and unifies multiple names for countries, companies, and universities by constructing databases with standard codes, processes the data to reduce distortion, and visualizes the results to facilitate accurate analysis.
The system enables accurate and intuitive data analysis by reducing distortion, allowing users to understand complex processes and assess trends effectively.
Smart Images

Figure KR2025018869_21052026_PF_FP_ABST
Abstract
Description
Device and method for processing, analyzing, and visualizing raw scientific and technological data
[0001] The present invention relates to an apparatus and method for processing, analyzing, and visualizing scientific and technological raw data that can reduce distortion and improve accuracy in the results of data analysis by standardizing and unifying multiple names for a country, multiple-named companies, and multiple-named universities during the processing of scientific and technological raw data (Raw Data File) including patent literature or papers.
[0002] In addition, the present invention sequentially uploads a raw data file of a specific technology downloaded from a data provider to a data processing module, unifies the multi-name country name, multi-name company name, and multi-name university name of the raw data file and writes them into a new first data processing file, uploads the first data processing file to a data analysis module to perform data indexing and analysis, writes the results into a second data processing file, and uploads the second data processing file to a data visualization module to perform visualization.
[0003] In addition, the present invention relates to a scientific and technological raw data processing, analysis, and visualization device that independently configures the data processing module, the data analysis module, and the visualization module, and outputs a first data processing file and a second data processing file at each stage of performing data processing, indexing, and visualization to analyze and resolve the cause of data distortion.
[0004] According to conventional technology, there are many disclosed cases of analyzing technological trends, etc., by utilizing raw data files downloaded from patent or paper database providers.
[0005] To analyze these technological trends, patent or paper data, which are the products of research and development from companies, universities, research institutions, and countries, are utilized. In particular, the results of the technological trend analysis are obtained through the number of patent disclosures, the number of published papers, and the number of citations in the patent or paper data.
[0006] Such patent or research paper data is vast, and the countries, companies, and universities around the world that disclosed it are also numerous; furthermore, they may be expressed differently depending on various factors such as region, language, and brand, and companies may have various names due to mergers, brand changes, or the use of abbreviations.
[0007] For example, “Samsung Electronics” is referred to as “Samsung Electronics” in English, and it may also appear under abbreviations such as “Samsung” or “SEC.” Handling such multi-name corporate names is crucial when analyzing data or integrating corporate information, and it can also be a significant variable in data distortion.
[0008] Furthermore, "Republic of Korea" is the official name, expressed as "Republic of Korea" in English, and abbreviations such as "Korea" or "KR" are used. This handling of multi-name country names is critical when analyzing data or integrating national information, and it becomes a significant variable that causes distortion in national data.
[0009] Furthermore, "Seoul National University" is referred to as "Seoul Dae" or "Seoul National University," and is also known by its abbreviation, "SNU." The "Massachusetts Institute of Technology" is sometimes abbreviated as "MIT." For consistent information processing in academic research and database management, it is necessary to identify and manage these multiple university names as a single entity.
[0010] In particular, using these various expressions or abbreviations in patents or papers is very commonplace.
[0011] Previously, it was difficult to perform accurate and reliable data analysis because raw patent or paper data was downloaded and analyzed as is. This is because patent applicants are diverse and expressed in a wide variety of ways.
[0012] Furthermore, since existing patent data-based data analysis primarily focuses on trend changes in absolute quantities by technology classification category, it is difficult to evaluate the relative changes in trend by category, intuitively understand them, and easily grasp the relative trend changes within each category.
[0013] In addition, it was difficult to update the analysis method because the analysis process was hard to understand.
[0014] Therefore, it is necessary to identify various names or abbreviations as a single country, company, or university and manage them consistently.
[0015] [Prior Art Literature]
[0016] [Patent Literature]
[0017] (Patent Document 1) KR 10-2543087B12023.06.08.
[0018] Therefore, the objective of the present invention is to reduce distortion in data analysis by constructing multi-name country names, multi-name company names, and multi-name university names of raw scientific and technological data, including patent or paper data, into a multi-name country name database (DB), a multi-name company name database (DB), and a multi-name university name database (DB), thereby unifying various abbreviations or expressions through consistent information processing.
[0019] In addition, another objective of the present invention is to provide a scientific and technological raw data processing, analysis, and visualization device and method that independently configures the data processing module, the data analysis module, and the visualization module, and outputs a first data processing file and a second data processing file at each step of performing data processing, indexing, and visualization to analyze and update the cause of data distortion.
[0020] To achieve the above objective, the present invention relates to a device for processing, analyzing, and visualizing scientific and technological raw data, including patents and papers, for reducing distortion of scientific and technological raw data and providing normalized analysis results, comprising: a database of multi-named country names (7) that integrates and manages multi-named country names including multiple names or expressions referring to the same country and assigns a standard code to a representative standard country name; a database of multi-named company names (6) that integrates and manages multi-named company names including multiple names or expressions referring to the same company and assigns a standard code to a representative standard company name; a database of multi-named university names (5) that integrates and manages multi-named university names including multiple names or expressions referring to the same university and assigns a standard code to a representative standard university name; and a database of company sales revenue (4) that manages the sales revenue and standard codes of the companies of the standard company names. The system is characterized by including: a data processing module (1) that is connected to each database (4, 5, 6, 7) and processes multiple names for a country, multiple names for companies, and multiple names for universities of a science and technology raw data file (10) to standardize and unify them, and outputs a first data processing file (20); a data analysis module (2) that uploads the first data processing file (20) output from the data processing module (1) and outputs a second data processing file (30) by counting and indexing the changes in the number of patents or papers published, the changes in the number of citations, and the changes in sales revenue; and a visualization analysis module (3) that uploads the second data processing file (30) output from the data analysis module (2) and visually expresses the growth trends of the index of changes in the number of patents or papers published, the index of changes in the number of patents or papers cited, and the index of changes in sales revenue.
[0021] The above visualization analysis module (3) is characterized by displaying growth trend lines of the patent or paper disclosure change index, patent or paper citation change index, and sales change index so that they can be visually compared with each other.
[0022] To achieve the above objective, the present invention relates to a method for processing, analyzing, and visualizing scientific and technological raw data, including patents and papers, for reducing distortion of scientific and technological raw data and providing normalized analysis results, comprising the steps of: constructing a database of multi-named country names (7) that integrates and manages multiple names or expressions referring to the same country and assigns a standard code to a representative standard country name (S11); constructing a database of multi-named company names (6) that integrates and manages multiple names or expressions referring to the same company and assigns a standard code to a representative standard company name (S12); constructing a database of multi-named university names (5) that integrates and manages multiple names or expressions referring to the same university and assigns a standard code to a representative standard university name (S13); and constructing a database of company sales revenue (4) that manages the sales revenue and standard codes of the companies with the standard company names (S14). A data processing step (S21, S22, S23, S24) for processing multiple names for a country, multi-named companies, and multi-named universities of a science and technology raw data file (10) based on a database (4, 5, 6, 7) to standardize and unify them; a step (S25) for outputting a first data processing file (20) processed in the data processing step (S20); an indexing step (S31, S32, S33, S34, S35) for counting and indexing the changes in the number of patents or papers published, the number of citations, and the sales revenue of the first data processing file; a step (S36, S37) for writing the count and indexing values into a second data processing file (30) and outputting the second data processing file (30);It is characterized by including a visualization step (S40) that visually expresses the growth trends of the fluctuation index of the number of patents or papers published, the fluctuation index of the number of patents or papers cited, and the fluctuation index of sales revenue of the second data processing file (30).
[0023] The data processing step (S20) comprises: a column creation step (S21) of creating a column (CODE) to record a data type such as a patent or paper, a column (CODE_NAME) to record a technology type, a column (COUNTRY) to record a country name, a column (NAME) to record a company or university name, a column (SCODE) to record a standard code of a company or university, a column (MARKET) to record sales information of a company, and a column (PCN) to record citation information in the first data processing file (20); and a step (S22) of writing data information including the data type of the patent or paper, country name, company name, citation information, and university name of the science and technology raw data file (10) into the columns of the first data processing file (20) in the plurality of columns created above. The method is characterized by including the step (S23) of entering the standard country, standard company, and standard university codes of the standard country, standard company, and standard university names among the data information into the columns of the first data processing file (20) using the multi-name university name database (5), multi-name company name database (6), and multi-name country name database (7); and the step (S24) of entering the sales revenue information of the company into the columns of the first data processing file (20) by comparing it with the standard code of the standard company using the company sales revenue database (4).
[0024] The indexing steps (S31, S32, S33, S34, S35) are characterized by including: a step of uploading the first data processing file (20) (S31); a step of counting the number of patent disclosures, papers, and citations by country, company, and university in the first data processing file (20) (S32); a step of calculating an R&D index by country, company, and university in which the number of patent disclosures in the first year of patent disclosure is set to 1 and the number of disclosures by subsequent years is expressed as a ratio for the number of patent disclosures counted by country, company, and university (S33); a step of calculating a technology innovation index by country, company, and university in which the number of patent or paper citations in the first year of patent or paper disclosure is set to 1 and the number of citations by subsequent years is expressed as a ratio for the number of patent or paper citations counted by country, company, and university (S34); and a step of calculating a market expansion index by which the company sales in the year of patent or paper disclosure is set to 1 and the sales by subsequent years are expressed as a ratio for the number of patent or paper citations counted by company (S35).
[0025] The above visualization step (S40) is characterized by including: a step (S41) of uploading the second data processing file (30); a step (S42) of visualizing the growth trend of the R&D index by country, company, and university entered in the second data processing file (30) using an exponential function; a step (S43) of visualizing the growth trend of the technology innovation index by country, company, and university entered in the second data processing file (30) using an exponential function; and a step (S44) of visualizing the growth trend of the market expansion index by company entered in the second data processing file (30) using an exponential function.
[0026] The present invention enables accurate data analysis results by significantly reducing data distortion through consistent information processing of multi-name country names, multi-name company names, and multi-name university names in raw scientific and technological data.
[0027] In addition, the present invention has the advantage of independently configuring the data processing module, the data analysis module, and the visualization module, and outputting a first data processing file and a second data processing file at each stage of sequentially performing data processing, indexing, and visualization, thereby enabling the analysis and correction of the causes of data distortion.
[0028] In addition, the present invention performs a systematic procedure for refining, analyzing, and visualizing raw scientific and technological data, allowing users to easily understand and use the complex data analysis process, and enables them to check the data refining and analysis content in the middle, so it can be beneficially utilized for non-experts or for data analysis education.
[0029] FIG. 1 is a configuration diagram (a) of an apparatus for processing, analyzing, and visualizing raw scientific and technological data files according to one embodiment of the present invention and a flowchart (b) of a method.
[0030] FIG. 2 is a configuration diagram of a data processing module (1) that uploads a science and technology raw data file and outputs a first data processing file according to an embodiment of the present invention.
[0031] FIG. 3 is a flowchart of the database construction step (S10) and data processing step (S20) performed by each database (4, 5, 6, 7) and data processing module (1).
[0032] FIG. 4 is a diagram of the configuration of a multi-name university name DB (5) for changing a multi-name university name (KEY) in a science and technology raw data file to a standard university name (VALUE) and assigning a standard code (SCODE) to one of the standard university names.
[0033] FIG. 5 is a diagram (a) of a multi-name company name DB (6) for changing a multi-name company name (KEY) of a science and technology raw data file to a standard company name (VALUE) and assigning a standard code (SCODE) to one of the standard company names, and a diagram (b) of a company sales revenue DB (4) for changing a multi-name company name (KEY) of a science and technology raw data file to a standard company name (VALUE), assigning a standard code (SCODE) to one of the standard company names, and assigning the company's sales revenue (MARKET) to the standard code.
[0034] FIG. 6 is a configuration diagram of a multi-name country name DB (7) for changing a multi-name country name (KEY) of a science and technology raw data file to a standard country name (VALUE) and assigning a standard code (SCODE) to one of the standard country names.
[0035] FIG. 7 is a configuration diagram of a data analysis module (2) that uploads a first data processing file, counts multiple data information, calculates multiple information indices, and outputs a second data processing file.
[0036] FIG. 8 is a flowchart of the indexing step (S30) and the visualization step (S40) performed sequentially by the data analysis module (2) and the visualization analysis module (3).
[0037] FIG. 9 is a diagram (a) illustrating the data format of information contained in a science and technology raw data file (10), and a diagram (b) illustrating the data format of a first data processing file (20) that is output as a new file by uploading the science and technology raw data file (10) to a data processing module (1) and processing the necessary information from the information of the science and technology raw data file (10).
[0038] FIG. 10 is a data format of a second data processing file (30) that is output by uploading a first data processing file (20) to a data analysis module (2), counting data information, and indexing it.
[0039] Figure 11 is an example of a graph illustrating growth trends by indexing the number of patent disclosures, number of papers, number of citations, and sales fluctuations.
[0040] FIG. 12 is an example of a UI (user interface) screen visualizing industrial technology trends, technology innovation index, market expansion index, and R&D investment index.
[0041] Hereinafter, embodiments of the present invention are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present invention. However, since it is clear that embodiments of the present invention can be implemented through various changes or modifications within the scope of the present invention, the present invention is not limited to the described embodiments. Furthermore, well-known components, circuits, functions, methods, and typical details regarding the embodiments of the present invention can be added and implemented by those skilled in the art, so they will not be described in detail.
[0042] Embodiments of the present invention may be implemented in a combined form of software and hardware, and the software and hardware forms may be described as components, modules, parts, etc., and may be implemented in the form of computer-readable program code implemented on a recording medium.
[0043] FIG. 1 illustrates a device for processing, analyzing, and visualizing a science and technology raw data file (10) according to the present invention and the sequence of operations of the device.
[0044] Referring to FIG. 1(a), the device for processing, analyzing, and visualizing a raw data file of science and technology according to the present invention includes a data processing module (1) that uploads a raw data file of science and technology (10) containing research and development information of industrial technologies such as artificial intelligence, robots, and Bitcoin, and processes a multi-name country name, a multi-name company name, and a multi-name university name to output a first data processing file (20); a data analysis module (2) that uploads the first data processing file (20), counts data information including the number of patent disclosures, the number of papers, the number of patent citations, and the number of paper citations, and indexes the value to output a second data processing file (30); and a visualization analysis module (3) that visually analyzes the trend fluctuations of the data information of the second data processing file (30).
[0045] Here, a corporate sales database (4), a multi-name university name database (5), a multi-name company name database (6), and a multi-name country name database (7) are also prepared for the operation of the data processing module (1).
[0046] Referring to FIG. 1(b), when the device has performed the database construction step (S10) and uploads the science and technology raw data file (10), it sequentially performs the data processing step (S20), the indexing step (S30), and the visualization step (S40).
[0047] Referring to FIG. 9(a), the science and technology raw data file (10) contains a lot of data information such as the applicant, country, publication date, and number of citations in the case of a patent data file, and in the case of a paper data file, it contains information such as the author's affiliation (institution), publication date, and country.
[0048] Upon closer examination, the applicant information contains numerous companies and universities engaged in research activities worldwide. It also includes a vast array of country names to which these companies and universities belong. These countries, universities, and companies are extensive and can be represented differently depending on various factors such as region, language, and brand; furthermore, companies may adopt various names due to mergers, brand changes, or the use of abbreviations.
[0049] The handling of such multiple names for country, company, and university names is very important when analyzing data or integrating corporate information, and it can also be a significant variable in data distortion. Here, multiple names refer to a single object existing under multiple names or expressions.
[0050] Accordingly, as illustrated in FIG. 1(a) and FIG. 2, in order to solve the problem of data distortion caused by multiple country names, multiple company names, and multiple university names, the data processing module (2) according to the present invention is connected to multiple university name DB (5), multiple company name DB (6), and multiple country name DB (7).
[0051] Referring to FIG. 4, the multi-name university name DB (5) constructs data of multi-name university names including multiple names or expressions referring to the same country in a column (KEY), constructs data of one representative standard university name corresponding to the multi-name university name in a column (VALUE), and assigns a university-specific standard code (SCODE) to the standard university name to normalize the multi-name university names.
[0052] Referring to FIG. 5(a), the multi-name company name DB (6) constructs data of multi-name company names that include multiple names or expressions referring to the same company in a column (KEY), constructs data of one standard company name corresponding to the multi-name company name in a column (VALUE), and provides a company-specific standard code (SCODE) to the standard company name to normalize the multi-name company names.
[0053] Referring to FIG. 6, the multi-name country name DB (8) constructs data of multi-name country names including multiple names or expressions referring to the same university in a column (KEY), constructs data of one standard country name corresponding to the multi-name country name in a column (VALUE), and provides a country-specific standard code (SCODE) to the standard country name to normalize the multi-name country names.
[0054] Meanwhile, as illustrated in FIG. 1(a), the present invention includes a corporate sales DB (4) for analyzing the market expansion trend of a company.
[0055] Referring to FIG. 5(b), the above-mentioned corporate sales DB (4) constructs multi-name corporate name data in a column (KEY), constructs standard corporate name data in a column (VALUE), constructs the standard code of the standard corporate name in a column (SCODE), and constructs the sales of the standard corporate name in a column (MARKET).
[0056] FIG. 2 illustrates a data processing module (1) that uploads a science and technology raw data file (10) according to one embodiment of the present invention, connects to the multi-name university name DB (5), multi-name company name DB (6) and multi-name country name DB (7), and outputs a first data processing file (20) that processes the university name, company name, and country name.
[0057] As illustrated in FIG. 9(b), the first data processing file (20) is composed of column (CODE) information representing patent (P) or paper (T) data information, column (CODE_NAME) information representing a specific type of technology, column (COUNTRY) information including a plurality of country names, column (NAME) information including companies and universities, column (SCODE) including a standard code, column (MARKET) containing company sales, and column (PCN) containing citation information of patent or paper data.
[0058] Referring to the flowchart of FIG. 3, the data processing module (1) according to the present invention generates a column (CODE) to record a data type such as a patent or paper, a column (CODE_NAME) to record a technology type, a column (COUNTRY) to record a country name, a column (NAME) to record a company or university name, a column (SCODE) to record a standard code of a company or university, a column (MARKET) to record sales information of a company, and a column (PCN) to record citation information in the first data processing file (20) (S21);
[0059] Data information such as the type of data, such as patents or papers, country name, company name, citation information, and university name of the science and technology raw data file (10) is entered into the columns of the first data processing file in the plurality of columns generated above (S22);
[0060] The above multi-name university name database (5), multi-name company name database (6) and multi-name country name database (7) are accessed, and the standard code assigned to the standard country, standard company, and standard university among the multi-name country name, multi-name company name, and multi-name university name data information is entered into the column of the first data processing file (20) (S23).
[0061] Next, the above corporate sales revenue database (4) is connected, and the corporate sales revenue information is entered into the columns of the first data processing file by comparing it with the standard code of the above standard corporate sales revenue (S24).
[0062] Finally, 1 data processing file (20) is output (S25).
[0063] As illustrated in FIG. 2, the reference code entry (S23) enters the company or university in the applicant (institution) column of the science and technology raw data file (10) of FIG. 9(a) into the column (NAME) of the first data processing file (20).
[0064] Referring to FIG. 4, the reference code entry (S23) uses the multi-name university name DB (5) to change the multi-name university name (KEY) of the science and technology raw data file (10) into a standard university name (VALUE) and assigns a reference code (SCODE) to one of the standard university names.
[0065] Referring to Fig. 5(a), the reference code entry (S23) uses the multi-name company name DB (6) to change the multi-name company name (KEY) into a standard company name (VALUE) and assigns a reference code (SCODE) to one of the company names.
[0066] Referring to FIG. 6, the reference code entry (S23) uses the multi-name country name DB (8) to change the multi-name country name (KEY) into a standard country name (VALUE) and assigns a reference code (SCODE) to one of the country names.
[0067] Referring to FIG. 5(b), the above-mentioned corporate sales entry (S24) compares the above-mentioned corporate standard code (SCODE) with the corporate standard code (SCODE) of the corporate sales DB (4), and if they match, the sales amount of the column (MARKET) of the corporate sales DB (4) is entered into the column (MARKET) of the first data processing file (20).
[0068] Through this process, the data processing module (2) outputs the first data processing file (20).
[0069] FIG. 7 illustrates a data analysis module (2) that uploads a first data processing file (20) to count data information including the number of patent disclosures, the number of papers, the number of patent citations, and the number of paper citations, and outputs a second data processing file (30) that calculates the patent disclosure fluctuation index, the number of papers fluctuation index, the citation fluctuation index, and the sales fluctuation index.
[0070] FIG. 8 illustrates a data analysis step (S30) performed by the data analysis module (2) and a visualization step (S40) subsequently performed by the visualization analysis module (3).
[0071] The above data analysis module (2) uploads the above first data processing file (20) (S31).
[0072] In the first data processing file (20) above, the number of patent disclosures, papers, and citations are counted by country, company, and university (S32).
[0073] For the number of patent disclosures counted for each of the above countries, companies, and universities, the number of patent disclosures in the first year of patent disclosure is set to 1, and the number of disclosures for subsequent years is expressed as a ratio to calculate the R&D index for each country, company, and university (S33).
[0074] For the number of patent or paper citations counted for each of the above countries, companies, and universities, an innovation index is calculated for each country, company, and university by setting the number of patent or paper publications in the first year of publication as 1 and expressing the number of publications in subsequent years as a ratio (S34).
[0075] A market expansion index is calculated by setting the company sales revenue of the year of patent or paper disclosure counted for each of the above companies to 1 and representing the sales revenue of subsequent years as a ratio (S35).
[0076] In this way, the count value and index value calculated by performing the indexing steps (S31, S32, S33, S34, S35) are entered into the second data processing file (30) (S36).
[0077] Next, output the second data processing file (30) (S37).
[0078] More specifically, the above-mentioned data information count (S32) uploads the first data processing file (20), counts the number of patent (P) or paper (T) data in the column (CODE) for the technical item column (CODE_NAME) in the first data processing file (20) of FIG. 9(b) to calculate the number of patent disclosures and papers, and calculates the number of citations by performing an addition operation on the numeric information of the column (PCN). In addition, it performs an addition operation on the sales revenue of a company using the company standard code (SCODE) of the column (NAME).
[0079] Referring to the second data processing file (30) of FIG. 10, the number of patent publications calculated by the data information count (S32) is stored in column (PAN), the number of papers is stored in column (TAN), the number of citations is stored in column (PCN), and the sales amount is stored in column (MARKET).
[0080] The above-mentioned index calculation (S33, S34, S35) calculates the sales revenue, number of patent disclosures, number of papers, and number of citations calculated from the above-mentioned data information count unit (32) as indices.
[0081] The indexing method for calculating these indices (S33, S34, S35) is as shown in the table below.
[0082] As shown in Table 1 below, let's assume the sales revenue of a certain company from 2000 to 2017. When the sales revenue from 2000 to 2017 is the value listed in the table below, the company's market expansion index can be calculated as follows.
[0083]
[0084] The above method of indexing sales revenue sets the sales of 100 won in 2000 as 1 and expresses subsequent sales as a ratio. This index conversion method is applied in the same way to the number of patent publications, papers, and citations.
[0085] When the second data processing file (30) with indexed values entered is uploaded (S41), the visualization analysis module (3) visualizes the growth trend of the R&D index by country, company, and university entered in the second data processing file (30) using an exponential function (S42), visualizes the growth trend of the technology innovation index by country, company, and university entered in the second data processing file (30) using an exponential function (S43), and visualizes the growth trend of the market expansion index by company entered in the second data processing file (30) using an exponential function (S44).
[0086] Figure 11 shows a graph illustrating the growth trend by indexing the number of patent disclosures, number of papers, number of citations, and fluctuations in sales revenue.
[0087] Through this, the company's sales growth trend can be intuitively grasped.
[0088] This indexing method allows for an intuitive understanding of growth trends by indexing changes in the number of patent disclosures, papers, and citations in the same way, in addition to sales revenue.
[0089] Figure 12 is a UI (User Interface) screen that visually analyzes industrial technology trends, technology innovation index, market expansion index, and R&D investment index using the number of patent disclosures, number of papers, number of citations, sales revenue, and R&D index, technology innovation index, and market expansion index.
[0090] Referring to FIG. 12, the UI screen above shows the growth trend of industrial technology using the values of the number of patent publications column (PAN) and the number of papers column (TAN) of the second data processing file (30).
[0091] This demonstrates that visual comparison is possible by simultaneously displaying the growth trend lines of the fluctuation index of the number of patent or paper publications, the fluctuation index of patent or paper citations, and the fluctuation index of sales revenue on a single coordinate system.
[0092] In this way, by resolving distortions caused by misrepresented country names, company names, and university names in raw scientific and technological data, it is possible to identify accurate and intuitive trends in industrial technology research and development and technological innovation through data analysis. Furthermore, it can be utilized very beneficially as it allows for assessing the potential for future market expansion.
[0093] [Explanation of the symbol]
[0094] 1 : Data Processing Module 2 : Data Analysis Module
[0095] 3: Visualization Analysis Module
[0096] 4 : Corporate Revenue Database 5 : Daqing University Name Database
[0097] 6 : Daqing Company Name Database 7 : Daqing Country Name Database
[0098] 10: Science and Technology Raw Data File 20: 1st Data Processing File
[0099] 30: 2nd Data Processing File
Claims
1. An apparatus for processing, analyzing, and visualizing scientific and technological raw data, including patents and papers, for reducing distortion of scientific and technological raw data and providing normalized analysis results, wherein A multi-name country name database (7) that constructs multi-name country names containing multiple names or expressions referring to the same country in a column (KEY), constructs one standard country name data corresponding to the multi-name country name in a column (VALUE), and assigns a country-specific standard code (SCODE) to the standard country name to normalize the multi-name country names; A database of multi-name company names (6) that includes multiple names or expressions referring to the same company, constructs a column (KEY) containing one standard company name data corresponding to the multi-name company name, constructs a column (VALUE) containing one standard company name data, and assigns a company-specific standard code (SCODE) to the standard company name to normalize the multi-name company names; A multi-name university name database (5) that constructs multi-name university names containing multiple names or expressions referring to the same university in a column (KEY), constructs one standard university name data corresponding to the multi-name university name in a column (VALUE), and assigns a university-specific standard code (SCODE) to the standard university name to normalize the multi-name university names; A corporate sales database (4) that manages standard corporate name standard codes and sales by constructing a multi-digit corporate name column (KEY), a standard corporate name column (VALUE), a standard corporate name standard code column (SCODE), and a standard corporate name sales column (MARKET); In order to standardize and unify the multiple names for a country, multi-named companies, and multi-named universities of the science and technology raw data file (10) by connecting to each database (4, 5, 6, 7), a column (CODE) to record the data type classified as patents and papers, a column (CODE_NAME) to record the technology type, a column (COUNTRY) to record the country name, a column (NAME) to record the company or university name, a column (SCODE) to record the standard code of the company or university, a column (MARKET) to record the company's sales information, and a column (PCN) to record the citation information, then the data type, technology type, country name, company name, university name, and citation information of the science and technology raw data file (10) are entered into the columns (CODE, CODE_NAME, COUNTRY, NAME, PCN), and the multiple names for a country, multi-named companies, and multi-named universities entered into the columns (COUNTRY, NAME) are used in the multiple names for a country database (7) and the multi-named companies A data processing module (1) that outputs a first data processing file (20) in which a standard country name, a standard company name, and a standard university name are changed using a database (6) and a multi-name university name database (5), a standard code assigned to the standard company name and the standard university name is entered into a column (SCODE), and the sales amount of the standard code (SCODE) of the company sales amount database (4) that matches the entered company standard code is entered into a column (MARKET); A data analysis module (2) that outputs a second data processing file (30) in which the first data processing file (20) output from the data processing module (1) is uploaded, and the number of patent disclosures, papers, and citations counted by country, company, and university, and the sales revenue calculated by addition by company are entered, and the R&D index by country, company, and university is set to 1 for the number of patent disclosures in the first year of patent disclosure and the number of disclosures by subsequent years is expressed as a ratio, the innovation index by country, company, and university is set to 1 for the number of citations in the first year of patent disclosure and the number of citations by subsequent years is expressed as a ratio, and the market expansion index by company is set to 1 for the company sales revenue in the year of patent disclosure and the sales revenue by subsequent years is expressed as a ratio; and A visualization analysis module (3) that uploads a second data processing file (30) output from the above data analysis module (2), visualizes and expresses the growth trends of the R&D index by country, company, and university, the innovation index by country, company, and university, and the market expansion index by company entered in the second data processing file (30) using an exponential function, and simultaneously displays the growth trend lines of the R&D index by country, company, and university, the innovation index by country, company, and university, and the market expansion index by company so that they can be visually compared with each other. A device for processing, analyzing, and visualizing scientific and technological raw data, characterized by including 2. A method for processing, analyzing, and visualizing raw scientific and technological data, wherein a data processing module (1) is connected to a multi-name country name database (7), a multi-name company name database (6), a multi-name university name database (5), and a company sales revenue database (4) to reduce distortion and normalize raw scientific and technological data including patents and papers, and then a data analysis module (2) and a visualization analysis module (3) provide analysis results. Step (S11) of constructing a database of multi-name country names in a database (7), wherein multi-name country names containing multiple names or expressions referring to the same country are constructed in a column (KEY), a single standard country name data corresponding to a multi-name country name is constructed in a column (VALUE), and a standard country name is assigned a country-specific standard code (SCODE) to the standard country name to construct a database of normalized multi-name country names; Step (S12) of constructing a database of multi-name corporate names that normalizes the multi-name corporate names by assigning a unique corporate standard code (SCODE) to the standard corporate name in a database of multi-name corporate names (6), constructing multi-name corporate names that include multiple names or expressions referring to the same corporate name in a column (KEY), constructing one standard corporate name data corresponding to the multi-name corporate name in a column (VALUE); Step (S13) of constructing a database of multi-name university names that normalizes the multi-name university names by assigning a university-specific standard code (SCODE) to the standard university name in a database of multi-name university names (5), constructing multi-name university names that include multiple names or expressions referring to the same university in a column (KEY), constructing one standard university name data corresponding to the multi-name university name in a column (VALUE); Step (S14) of constructing a database that manages the standard code and sales of a standard company name by constructing a multi-identified company name column (KEY), a standard company name column (VALUE), a standard code column (SCODE) of the standard company name and a sales of the standard company name column (MARKET) in the company sales database (4); A data processing step (S20) in which, in the data processing module (1), a first data processing file (20) is generated and output by processing the multiple names for a country, multi-named companies, and multi-named universities of the science and technology raw data file (10) based on the database (4, 5, 6, 7) to standardize and unify them; A data analysis step (S30) in which a second data processing file (30) is generated and output in the data analysis module (2), wherein the values of the changes in the number of patents or papers published, the changes in the number of citations, and the changes in sales revenue of the first data processing file (20) are counted and indexed. The visualization analysis module (3) includes a visualization step (S40) that visually represents the growth trends of the patent or paper disclosure change index, patent or paper citation change index, and sales change index of the second data processing file (30). The above data processing step (S20) is A column generation step (S21) for generating a column (CODE) to record data types classified as patents and papers, a column (CODE_NAME) to record technology types, a column (COUNTRY) to record country names, a column (NAME) to record company or university names, a column (SCODE) to record standard codes of companies or universities, a column (MARKET) to record sales information of companies, and a column (PCN) to record citation information in the first data processing file (20); Step (S22) of entering data type, technology type, country name, company name, university name, and citation information of the science and technology raw data file (10) into the columns (CODE, CODE_NAME, COUNTRY, NAME, PCN) of the first data processing file (20); A step (S23) of changing the multi-name country name, multi-name company name, and multi-name university name entered in the column (COUNTRY, NAME) into a standard country name, standard company name, and standard university name using the multi-name country name database (7), multi-name company name database (6), and multi-name university name database (5), and entering the standard code assigned to the standard company name and standard university name into the column (SCODE); Step (S24) of entering the sales amount of the reference code (SCODE) of the corporate sales amount database (4) that matches the entered corporate reference code into a column (MARKET); and It includes the step (S25) of outputting the first data processing file (20), and The above data analysis step (S30) is Step (S31) of uploading the first data processing file (20) above; A step (S32) of counting the number of patent disclosures, papers, and citations by country, company, and university in the first data processing file (20), and performing an addition operation on the sales revenue by company; A step (S33) of calculating an R&D index for each country, company, and university, wherein the number of patent disclosures in the first year of patent disclosure is set to 1 and the number of disclosures in subsequent years is expressed as a ratio; A step (S34) of calculating an innovation index for each country, company, and university, wherein the number of patent or paper citations counted by country, company, and university is set to 1 for the year of initial publication of the patent or paper and the number of subsequent yearly citations is expressed as a ratio; A step (S35) of calculating a market expansion index in which the company's sales revenue for the year of patent or paper disclosure counted by company is set to 1, and the subsequent yearly sales revenue is expressed as a ratio; and The method includes the steps (S36, S37) of entering the count values of the number of patent disclosures, the number of papers and citations, the sales addition operation value and the indexed value into the second data processing file (30), and outputting the second data processing file (30). The above visualization step (S40) Step (S41) in which the above second data processing file (30) is uploaded; A step (S42) of visualizing the growth trend of R&D indices by country, company, and university entered in the second data processing file (30) as an exponential function; A step (S43) of visualizing the growth trend of innovation indices by country, company, and university entered in the second data processing file (30) as an exponential function; and The method includes a step (S44) of visualizing the growth trend of the market expansion index by company entered in the second data processing file (30) as an exponential function, and simultaneously displaying the growth trend lines of the R&D index, innovation index, and market expansion index so that they can be visually compared with each other. A method for processing, analyzing, and visualizing raw scientific and technological data characterized by