Document Search System

The document search system addresses the challenge of accurately searching intellectual property documents by using a processing unit to generate weight and thesaurus dictionary data, enabling high-accuracy searches with a simple input method.

JP7700037B2Active Publication Date: 2025-06-30SEMICON ENERGY LAB CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021515317
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-04-26
Filing Date
2020-04-16
Publication Date
2025-06-30
Estimated Expiration
2040-04-16

AI Technical Summary

Technical Problem

Existing document search systems require high user skills to search patent documents accurately, and there is a need for a system that can search documents related to intellectual property with high accuracy using a simple input method.

Method used

A document search system comprising an input unit, a database, a storage unit, and a processing unit that generates weight dictionary data and thesaurus dictionary data based on reference document data, allowing for the extraction of search words and the generation of search data for accurate document retrieval.

Benefits of technology

The system enables high-accuracy document search, particularly for intellectual property-related documents, with a simple input method, reducing user burden and minimizing search result variability due to user skill differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700037000002
    Figure 0007700037000002
  • Figure 0007700037000003
    Figure 0007700037000003
  • Figure 0007700037000004
    Figure 0007700037000004
Patent Text Reader

Abstract

This invention realizes a highly accurate document search, particularly a search for documents pertaining to intellectual property, with a simple input method. A processing unit includes: a function for generating sentence analysis data from sentence data input to an input unit; a function for extracting a search word from among words included in the sentence analysis data; and a function for generating first search data from the search word, on the basis of weighting dictionary data and synonym dictionary data. A storage unit stores second search data, which is generated by correction of the first search data by a user. The processing unit updates the synonym dictionary data in accordance with the second search data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to a document search system and a document search method.

[0002] Note that one aspect of the present invention is not limited to the above technical field. Examples of the technical field of one aspect of the present invention include semiconductor devices, display devices, light-emitting devices, power storage devices, memory devices, electronic devices, lighting devices, input devices (for example, touch sensors, etc.), input / output devices (for example, touch panels, etc.), their driving methods, or their manufacturing methods.

Background Art

[0003] By conducting a prior art search on an invention before filing, it is possible to investigate whether there are any related intellectual property rights. Patent documents and papers at home and abroad obtained by conducting a prior art search can be used to confirm the novelty and inventiveness of the invention and to determine whether to file a patent. In addition, by conducting an invalidation search of patent documents, it is possible to investigate whether there is a risk of invalidating one's own patent rights or whether it is possible to invalidate the patent rights owned by others.

[0004] For example, in a system for searching patent documents, when a user inputs a keyword, patent documents including the keyword can be output.

[0005] In order to conduct a prior art search with high accuracy using such a system, high skills are required of the user, such as searching with appropriate keywords and extracting necessary patent documents from many output patent documents.

[0006] In addition, in various applications, the use of artificial intelligence is being considered. In particular, it is expected that by using artificial neural networks or the like, a computer with higher performance than a conventional Neumann-type computer can be realized, and in recent years, various studies on constructing artificial neural networks on electronic circuits have been advanced.

[0007] For example, Patent Document 1 discloses an invention in which a storage device using a transistor having an oxide semiconductor in a channel formation region holds weight data necessary for calculations using an artificial neural network.

Prior Art Documents

Patent Documents

[0008]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0009] Therefore, one aspect of the present invention is to provide a document search system capable of searching documents with high accuracy. Or, one aspect of the present invention is to provide a document search method capable of searching documents with high accuracy. Or, one aspect of the present invention is to realize document search with high accuracy, particularly search for documents related to intellectual property, with a simple input method.

[0010] The description of a plurality of problems does not prevent the existence of each other's problems. One form of the present invention does not need to solve all the exemplified problems. Also, problems other than those listed will become apparent from the description of this specification, and such problems can also be problems of one form of the present invention.

Means for Solving the Problems

[0011] One aspect of the present invention has an input unit, a database, a storage unit, and a processing unit. The database has a function of storing a plurality of reference document data, a weight dictionary data, and a thesaurus dictionary data. The processing unit has a function of generating the weight dictionary data and the thesaurus dictionary data based on the reference document data, a function of generating document analysis data from the text data input to the input unit, a function of extracting a search word from the words included in the document analysis data, and a function of generating first search data from the search word based on the weight dictionary data and the thesaurus dictionary data. The storage unit has a function of storing second search data generated by the user modifying the first search data. The processing unit has a function of updating the thesaurus dictionary data according to the second search data. This is a document search system.

[0012] In one aspect of the present invention, it is preferable that the processing unit has a function of generating reference document analysis data from the reference document data, and a function of extracting a plurality of keywords and related words corresponding to the keywords from the words included in the reference document analysis data. This is a document search system.

[0013] In one aspect of the present invention, it is preferable that the weight dictionary data is data generated by extracting the frequency of occurrence of keywords from the words included in the reference document analysis data and assigning a first weight corresponding to the frequency of occurrence to each of the keywords. This is a document search system.

[0014] In one aspect of the present invention, it is preferable that the first weight is a value based on the inverse document frequency of the keyword in the reference document analysis data. This is a document search system.

[0015] In one aspect of the present invention, it is preferable that the thesaurus dictionary data is data generated by assigning a second weight to each of the related words. This is a document search system.

[0016] In one aspect of the present invention, a document retrieval system is preferred, in which the second weight is the product of a value based on the similarity or distance between the distributed representation vector of a related word and the distributed representation vector of a keyword, and the first weight of the keyword.

[0017] In one aspect of the present invention, a document retrieval system is preferred, in which the distributed representation vector is a vector generated using a neural network.

[0018] In one aspect of the present invention, a document retrieval system is preferred, in which the processing unit includes transistors, and the transistors have metal oxides in the channel formation regions.

[0019] In one aspect of the present invention, a document retrieval system is preferred, in which the processing unit includes transistors, and the transistors have silicon in the channel formation regions.

[0020] One aspect of the present invention is a document retrieval method, which includes generating weight dictionary data and synonym dictionary data based on a plurality of reference document data, generating document analysis data from document data, extracting search words from the words included in the document analysis data, generating first search data from the search words based on the weight dictionary data and the synonym dictionary data, updating the synonym dictionary data according to second search data generated by the user's modification of the first search data, assigning scores to the reference document data based on the second search data, and generating ranking data by ranking the plurality of reference document data based on the scores.

[0021] In one aspect of the present invention, a document retrieval method is preferred, which includes generating reference document analysis data from the reference document data, and extracting a plurality of keywords and related words of the keywords from the words included in the reference document analysis data.

[0022] In one aspect of the present invention, a weight dictionary data is data generated by extracting the occurrence frequency of keywords from among the words included in the reference document analysis data and assigning a first weight according to the occurrence frequency to each of a plurality of keywords. A document search method is preferable.

[0023] In one aspect of the present invention, a document search method is preferable, in which the first weight is a value based on the inverse document frequency of the keyword in the reference document analysis data.

[0024] In one aspect of the present invention, a document search method is preferable, in which a thesaurus dictionary data is data generated by assigning a second weight to each of related words.

[0025] In one aspect of the present invention, a document search method is preferable, in which the second weight is a product of a value based on the similarity or distance between the distributed representation vector of the related word and the distributed representation vector of the keyword, and the first weight of the keyword.

[0026] In one aspect of the present invention, a document search method is preferable, in which the distributed representation vector is a vector generated using a neural network.

[0027] For other aspects of the present invention, they are described in the embodiments described below and in the drawings.

Advantages of the Invention

[0028] According to one aspect of the present invention, a document search system capable of searching documents with high accuracy can be provided. Alternatively, according to one aspect of the present invention, a document search method capable of searching documents with high accuracy can be provided. Alternatively, according to one aspect of the present invention, a simple input method can be used to realize highly accurate document search, particularly for documents related to intellectual property.

[0029] The description of multiple effects does not prevent the existence of other effects. Also, one embodiment of the present invention does not necessarily have all of the illustrated effects. Further, regarding one embodiment of the present invention, other problems, effects, and novel features will become apparent from the description and drawings of this specification.

Brief Description of the Drawings

[0030] FIG. 1 is a block diagram showing an example of a document search system. FIG. 2 is a flowchart for explaining a document search method. FIG. 3 is a flowchart for explaining a document search method. FIG. 4 is a flowchart for explaining a document search method. FIG. 5 is a flowchart for explaining a document search method. FIGS. 6A to 6C are schematic diagrams for explaining a document search method. FIG. 7 is a schematic diagram for explaining a document search method. FIG. 8 is a schematic diagram for explaining a document search method. FIG. 9 is a schematic diagram for explaining a document search method. FIG. 10 is a flowchart for explaining a document search method. FIG. 11 is a flowchart for explaining a document search method. FIG. 12 is a flowchart for explaining a document search method. FIGS. 13A and 13B are diagrams showing a configuration example of a neural network. FIG. 14 is a diagram showing a configuration example of a semiconductor device. FIG. 15 is a diagram showing a configuration example of a memory cell. FIG. 16 is a diagram showing a configuration example of an offset circuit. FIG. 17 is a timing chart.

Embodiments for Carrying Out the Invention

[0031] The embodiments of the present invention will be described below. However, one embodiment of the present invention is not limited to the following description, and it is easily understood by those skilled in the art that the form and details can be variously changed without departing from the gist and scope of the present invention. Therefore, one embodiment of the present invention is not construed as being limited to the description of the embodiments shown below.

[0032] In this specification and the like, ordinal numbers such as "first", "second", and "third" are attached to avoid confusion of components. Therefore, they do not limit the number of components. Also, they do not limit the order of components. For example, in one of the embodiments of this specification and the like, the component referred to as "first" may be the component referred to as "second" in other embodiments or in the claims. Also, for example, the component referred to as "first" in one of the embodiments of this specification and the like may be omitted in other embodiments or in the claims.

[0033] In the drawings, the same reference numerals may be given to the same elements, elements having the same or similar functions, elements of the same material, or elements formed simultaneously, and the repeated description may be omitted.

[0034] In this specification, for example, the power supply potential VDD may be described by omitting it as the potential VDD, VDD, etc. The same applies to other components (for example, signals, voltages, circuits, elements, electrodes, wirings, etc.).

[0035] Also, when the same reference numeral is used for a plurality of elements, particularly when it is necessary to distinguish them, an identification symbol such as "_1", "_2", "[n]", "[m,n]" may be appended to the reference numeral for description. For example, the second wiring GL is described as wiring GL[2].

[0036] (Embodiment 1) In this embodiment, a document search system and a document search method according to one aspect of the present invention will be described with reference to FIGS. 1 to 12.

[0037] In this embodiment, as an example of a document retrieval system, a document retrieval system that can be used for intellectual property retrieval will be described. Note that the document retrieval system according to one aspect of the present invention is not limited to the use of intellectual property retrieval, and can also be used for retrieval other than intellectual property.

[0038] FIG. 1 shows a block diagram of a document retrieval system 10. The document retrieval system 10 includes an input unit 20, a processing unit 30, a storage unit 40, a database 50, an output unit 60, and a transmission path 70.

[0039] Data (such as text data 21) is supplied to the input unit 20 from outside the document retrieval system 10. Further, data (such as retrieval data 61) output from the output unit 60 is supplied to the input unit as modified data (such as retrieval data 62) generated by a user who modifies the data using the document retrieval system. The text data 21 and the retrieval data 62 are supplied to the processing unit 30, the storage unit 40, or the database 50 via the transmission path 70.

[0040] In this specification and the like, data of documents related to intellectual property is referred to as document data. The above text data is data corresponding to a part of the document data. Specifically, examples of the document data include data of publications such as patent documents (published patent gazettes, patent gazettes, etc.), utility model gazettes, design gazettes, and papers. It is not limited to publications issued in Japan, and publications issued in various countries around the world can be used as document data related to intellectual property. Note that the document data corresponds to data referred to for text data including the text to be retrieved. Therefore, the document data may sometimes be referred to as reference document data.

[0041] The above text data 21 is part of the above reference document data. Specifically, the specification, claims, and abstract included in the patent document can each be used in part or in whole as the text data 21. For example, the forms, examples, or claims for implementing a specific invention may be used as the text data 21. Similarly, for the text included in other publications such as papers, part or all of it can be used as the text data 21.

[0042] Documents related to intellectual property are not limited to publications. For example, document files independently owned by users or user groups of the document search system can also be used as the text data 21.

[0043] Furthermore, examples of documents related to intellectual property include articles explaining inventions, utility models, or designs, or industrial products.

[0044] The text data 21 can have, for example, patent documents of a specific applicant or patent documents in a specific technical field.

[0045] The text data 21 can have not only descriptions of intellectual property itself (such as specifications), but also various information related to the intellectual property (such as bibliographic information). Examples of such information include the applicant of the patent, technical field, application number, publication number, status (pending, registered, withdrawn, etc.).

[0046] The text data 21 preferably has date information related to intellectual property. Examples of date information include, for example, the application date, publication date, registration date, etc. if the intellectual property is a patent document, and the release date, etc. if the intellectual property is technical information of an industrial product.

[0047] In this way, since the text data 21 has various information related to intellectual property, various search ranges can be selected using the document search system.

[0048] The processing unit 30 has a function of performing operations, inferences, etc. using data supplied from the input unit 20, the storage unit 40, the database 50, etc. The processing unit 30 can supply operation results, inference results, etc. to the storage unit 40, the database 50, the output unit 60, etc.

[0049] It is preferable to use a transistor having a metal oxide in the channel formation region for the processing unit 30. Since the off-current of the transistor is extremely small, by using the transistor as a switch for holding the charge (data) flowing into the capacitive element that functions as a memory element, the data holding period can be ensured over a long term. By using this characteristic for at least one of the register and the cache memory that the processing unit 30 has, the processing unit 30 can be operated only when necessary, and in other cases, the processing unit 30 can be turned off by saving the information of the previous processing in the memory element. That is, normally-off computing becomes possible, and power consumption of the document search system can be reduced.

[0050] In this specification and the like, a transistor using an oxide semiconductor or a metal oxide in the channel formation region is called an Oxide Semiconductor transistor, or an OS transistor for short. The channel formation region of the OS transistor preferably has a metal oxide.

[0051] In this specification and the like, a metal oxide is an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also referred to as Oxide Semiconductor or simply OS), etc. For example, when a metal oxide is used for the semiconductor layer of a transistor, the metal oxide may be referred to as an oxide semiconductor. That is, when a metal oxide has at least one of an amplification action, a rectification action, and a switching action, the metal oxide can be called a metal oxide semiconductor, abbreviated as OS.

[0052] The metal oxide included in the channel formation region preferably contains indium (In). When the metal oxide included in the channel formation region is a metal oxide containing indium, the carrier mobility (electron mobility) of the OS transistor increases. Further, the metal oxide included in the channel formation region is preferably an oxide semiconductor containing element M. Element M is preferably aluminum (Al), gallium (Ga), tin (Sn), or the like. Applicable elements for other element Ms include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), tungsten (W), and the like. However, there may be cases where a plurality of the aforementioned elements are combined as element M. Element M is, for example, an element having a high binding energy with oxygen. For example, it is an element having a higher binding energy with oxygen than indium. Further, the metal oxide included in the channel formation region is preferably a metal oxide containing zinc (Zn). The metal oxide containing zinc may be likely to crystallize.

[0053] The metal oxide included in the channel formation region is not limited to a metal oxide containing indium. The semiconductor layer may be, for example, a metal oxide containing zinc and not containing indium, such as zinc tin oxide or gallium tin oxide, a metal oxide containing gallium, or a metal oxide containing tin.

[0054] Further, a transistor including silicon in the channel formation region may be used in the processing unit 30.

[0055] Further, it is preferable to use, in combination, a transistor including an oxide semiconductor in the channel formation region and a transistor including silicon in the channel formation region in the processing unit 30.

[0056] The processing unit 30 includes, for example, an arithmetic circuit or a central processing unit (CPU).

[0057] The processing unit 30 may have a microprocessor such as a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The microprocessor may be configured to be implemented by a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array) or an FPAA (Field Programmable Analog Array). The processing unit 30 can perform various data processing and program controls by interpreting and executing instructions from various programs using a processor. Programs executable by the processor are stored in at least one of the memory area of the processor and the storage unit 40.

[0058] The processing unit 30 may have a main memory. The main memory has at least one of a volatile memory such as a RAM (Random Access Memory) and a non-volatile memory such as a ROM (Read Only Memory).

[0059] As the RAM, for example, DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), etc. are used, and a memory space is virtually allocated and used as the working space of the processing unit 30. The operating system, application programs, program modules, program data, look-up tables, etc. stored in the storage unit 40 are loaded into the RAM for execution. These data, programs, and program modules loaded into the RAM are directly accessed and operated on by the processing unit 30, respectively.

[0060] The ROM can store the BIOS (Basic Input / Output System), firmware, etc. that do not require rewriting. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), etc. Examples of EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory) that enables erasure of stored data by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, etc.

[0061] The storage unit 40 has a function of storing the programs executed by the processing unit 30. Further, the storage unit 40 may have a function of storing the calculation results and inference results generated by the processing unit 30, as well as the data input to the input unit 20. The storage unit 40 also has a function of storing the search data 62 input to the input unit 20 as search data 41 in the storage unit 40. The search data 41 stored in the storage unit 40 is used to update the synonym dictionary data described later.

[0062] The storage unit 40 has at least one of a volatile memory and a non-volatile memory. The storage unit 40 may have, for example, a volatile memory such as DRAM or SRAM. The storage unit 40 may have, for example, a non-volatile memory such as ReRAM (Resistive Random Access Memory, also called a resistive change type memory), PRAM (Phase change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory, also called a magnetic resistance type memory), or a flash memory. Further, the storage unit 40 may have a recording media drive such as a hard disc drive (HDD) and a solid state drive (SSD).

[0063] The database 50 has at least a function of storing reference document data 51 to be searched, weight dictionary data 52, and similar word search data 53. Further, the database 50 may have a function of storing the calculation results and inference results generated by the processing unit 30, and the data input to the input unit 20. Note that the storage unit 40 and the database 50 may not be separated from each other. For example, the document search system 10 may have a storage unit having the functions of both the storage unit 40 and the database 50.

[0064] The reference document data 51 is data of a plurality of documents related to intellectual property. The weight dictionary data 52 is data generated by extracting the occurrence frequencies of a plurality of keywords from the words included in the reference text analysis data obtained by analyzing the reference document data 51 and assigning weights corresponding to the occurrence frequencies to each of the plurality of keywords. The similar word search data 53 is data generated by extracting related words corresponding to the keywords from the words included in the reference text analysis data and assigning weights corresponding to the similarity degrees to each of the related words.

[0065] In addition, the database 50 has a function of storing inverse document frequency (hereinafter referred to as IDF) data (hereinafter referred to as IDF data) necessary for generating the weight dictionary data 52 and the similar word search data 53. IDF represents the difficulty of a certain word appearing in a document. The IDF of a word that appears in many documents is small, and the IDF of a word that appears only in some documents is high. Therefore, it can be said that a word with a high IDF is a characteristic word in the reference text analysis data. It is preferable to use the IDF data for calculating the appearance frequency of the above keywords.

[0066] In addition, the extraction of search words from the text data can also be performed based on the IDF. For example, words with an IDF equal to or greater than a certain value may be extracted as search words, or any number of words may be extracted as search words in descending order of IDF.

[0067] In addition, the database 50 has a function of storing vector data necessary for calculating related words corresponding to the keywords. The related words are extracted from the words included in the reference text analysis data based on the degree of similarity or the proximity of the distributed representation vector of the word and the distributed representation vector of the keyword. For calculating the weight of the related word, it is preferable to use the product of a value based on the similarity or distance between the distributed representation vector of the related word and the distributed representation vector of the keyword and the weight of the keyword. Alternatively, for calculating the weight of the related word, a value based on the similarity or distance between the distributed representation vector of the related word and the distributed representation vector of the keyword may be used. The search accuracy can be improved by setting the weight of the related word based on both the similarity between the related word and the keyword and the weight of the keyword itself. Examples of related words include synonyms, similar words, antonyms, hypernyms, hyponyms, etc.

[0068] The above search data 61 corresponds to data generated by extracting search words included in the text data 21 and referring to the thesaurus data and the weight dictionary data. The search data is data in which a keyword corresponding to the search word and a related word corresponding to the keyword are each given a weight. By each of the keyword and the related word having a weight, a score based on the weight can be given to the reference document data in which the keyword or the related word hits. The search data 62 corresponds to data in which the above weight is corrected by a user operation in the search data 61.

[0069] The output unit 60 has a function of supplying search data outside the document search system 10. For example, the search data generated in the processing unit 30 can be supplied to a display device or the like provided outside the document search system 10. The user can confirm the search data generated via a display device or the like provided outside the document search system 10.

[0070] The transmission path 70 has a function of transmitting data. Transmission and reception of data between the input unit 20, the processing unit 30, the storage unit 40, the database 50, and the output unit 60 can be performed via the transmission path 70.

[0071] FIG. 2 is a diagram showing a flow for explaining a document search method using the document search system 10 described in FIG. 1.

[0072] In the flow illustrated in FIG. 2, first, registration of reference document data is performed on the database 50 (step S11). This step of performing the registration may be configured to be performed during subsequent steps.

[0073] Next, creation of the weight dictionary data is performed (step S12). The weight dictionary data creation flow in this step S12 will be described with reference to FIG. 3 described later.

[0074] Next, synonym dictionary data is created (step S13). The synonym dictionary data creation flow in this step S13 will be described with reference to FIG. 4 described later. Note that step S13 may be performed by swapping with step S12, or may be performed at the same timing.

[0075] Next, document data is input (step S14). This input of document data is performed via a graphical user interface (GUI) such as a display device provided outside the document search system 10.

[0076] Next, search words are extracted from the document data (step S15). The search word extraction flow in this step S15 will be described with reference to FIG. 5 described later.

[0077] Next, search data is created (step S16). The creation of search data is performed with reference to the search words, the weight dictionary data, and the similar word dictionary data. The search data in this step S16 will be described with reference to FIG. 7 and the like described later.

[0078] Next, the search data based on the search data is displayed (step S17). The display is performed by outputting the search data to a GUI such as a display device provided outside the document search system 10.

[0079] Next, the search data displayed in step S17 is corrected (step S18). This correction is performed by the user correcting the value of the weight data included in the search data displayed on the display device provided outside the document search system 10.

[0080] Next, a search is executed based on the corrected search data (step S19). The search execution flow in this step S19 will be described with reference to FIG. 11 described later.

[0081] The search data corrected in step S18 is stored in a storage unit or the like (step S20).

[0082] After the search is executed in step S19, it is determined whether to end the search (step S21). If continuing, return to step S14 and input the document data again. If ending, the search ends.

[0083] After saving the search data corrected in step S20, the synonym dictionary data is updated (step S22). That is, the data created in the creation of the synonym dictionary data shown in step S13 is updated. The update flow of the synonym dictionary data in this step S22 will be described with reference to FIG. 10 and the like described later.

[0084] According to the flowchart of FIG. 2, in the document search method of one aspect of the present invention, the synonym dictionary data can be updated using the search data corrected by the user. Therefore, a document search method capable of searching documents with high accuracy can be provided. Alternatively, with a simple input method, high-accuracy document search, particularly for documents related to intellectual property, can be realized.

[0085] FIG. 3 is a diagram showing a flow for generating the weight dictionary data shown in step S12 described in FIG. 2.

[0086] First, a plurality of reference document data (hereinafter referred to as document data TD REF ) is input to the processing unit 30 via the input unit 20 (step S41). Step S41 corresponds to step S11 described above.

[0087] Next, word segmentation processing is performed on the document data TD REF . Thereafter, it is preferable to perform processing for correcting unnecessary word segmentation processing.

[0088] Next, morphological analysis is performed on the document data TD REF on which the word segmentation processing has been performed (step S43).

[0089] Next, for the data on which morphological analysis has been performed, document analysis data AD REFGenerate (reference text analysis data) (step S44). In morphological analysis, a sentence written in natural language can be divided into morphemes (the smallest units that have meaning as a language), and the part of speech of the morphemes can be determined. As a result, for example, document data TD REF from which only nouns are extracted can be used as text analysis data AD REF .

[0090] Next, for the text analysis data AD REF , calculate the IDF of the words included in the text analysis data AD REF and generate IDF data ID (step S45). The IDF data ID includes a word (Word) and a normalized IDF. The IDF data ID includes a word (Word) that becomes a keyword and a normalized IDF.

[0091] The IDF(t) of a certain word t is obtained by normalizing the idf(t) in formula (1). The method of normalization is not particularly limited. For example, the idf(t) can be normalized by formula (2). In formula (1), N is the total number of documents (the number of reference text analysis data AD ref ), and df(t) is the number of documents in which a certain word t appears (the number of reference text analysis data AD ref ). In formula (2), idf MAX is the maximum value of the idf(t) of the words included in the reference text analysis data AD ref , and idf MIN is the minimum value of the idf(t) of the words included in the reference text analysis data AD ref .

[0092]

Equation

[0093] Words with high IDF are the text analysis data AD REFIt can be said that these are characteristic words that are difficult to appear. Therefore, by estimating the IDF data ID standardized for each word, it is possible to extract keywords, which are characteristic words for retrieving a desired document, and the standardized IDF.

[0094] Next, in the IDF data ID, the IDF attached to each keyword is used as weight data, and weight dictionary data with weight data attached to each keyword is generated (step S46). As described above, it can be said that words with a high IDF are characteristic words in the reference document analysis data. By extracting the IDF, the frequency of appearance for each keyword can be estimated, and weight dictionary data in which weight data corresponding to the frequency of appearance is associated with each keyword can be generated. The generated weight dictionary data can be stored in the database 50.

[0095] According to the flowchart of FIG. 3, weight dictionary data can be generated based on the reference document data stored in the database. For each characteristic word (keyword) in the document data, the importance (weight) of each keyword can be estimated by estimating with a numerical value standardized by the IDF. Therefore, it is possible to provide a document retrieval method capable of retrieving documents with high accuracy. Alternatively, with a simple input method, it is possible to realize highly accurate document retrieval, particularly retrieval of documents related to intellectual property.

[0096] FIG. 4 is a diagram showing a flow for generating the synonym dictionary data shown in step S12 described in FIG. 2.

[0097] First, the document data TD REF is input to the processing unit 30 via the input unit 20 (step S51). Step S51 corresponds to step S11 described above. Note that step S51 corresponds to the same process as step S41.

[0098] Next, the document data TD REFPerform word segmentation processing (step S52). Thereafter, it is preferable to perform processing for correcting unnecessary word segmentation processing. Note that step S52 corresponds to the same processing as step S42.

[0099] Next, for the document data TD REF for which word segmentation processing has been performed, perform morphological analysis (step S53). Note that step S53 corresponds to the same processing as step S43.

[0100] Next, for the data for which morphological analysis has been performed, generate document analysis data AD REF (reference document analysis data) (step S54). Note that step S54 corresponds to the same processing as step S44.

[0101] Next, for the document analysis data AD REF calculate the IDF of the words included in the document analysis data AD REF and generate IDF data ID (step S55). Note that step S55 corresponds to the same processing as step S45. By estimating the IDF data ID normalized for each word, it is possible to extract keywords, which are characteristic words for searching for a desired document, and the normalized IDF.

[0102] Next, for the document analysis data AD REF extract the words included in the data, generate a distributed representation vector for each word, and generate vector data VD (step S56).

[0103] The distributed representation of a word is also called word embedding. The distributed representation vector of a word is a vector that represents a word as a continuous value quantified for each feature element (dimension). Words with similar meanings have vectors that are also close.

[0104] The processing unit 30 preferably generates a distributed representation vector of words using a neural network. The learning of the neural network is performed by supervised learning. Specifically, a certain word is given to the input layer, and the surrounding words of the word are given to the output layer, and the neural network is made to learn the probability of the surrounding words for a certain word. The intermediate layer (hidden layer) preferably has a relatively low-dimensional vector of 10 dimensions or more and 1000 dimensions or less. The vector after learning is the distributed representation vector of the word.

[0105] The distributed representation of words can be performed, for example, using the open-sourced algorithm Word2vec. Word2vec vectorizes words including their features and semantic structures based on the hypothesis that words used in the same context have the same meaning.

[0106] In the vectorization of words, by generating the distributed representation vector of words, it is possible to calculate the similarity and distance between words by operations between vectors. When the similarity between two vectors is high, it can be said that the relationship between the two vectors is high. Also, when the distance between two vectors is close, it can be said that the relationship between the two vectors is high.

[0107] Also, while the one-hot representation assigns one dimension to one word, in the distributed representation, since a word can be represented by a low-dimensional real-valued vector, it can be represented with a small number of dimensions even when the vocabulary size increases. Therefore, even if the number of words included in the corpus is large, the amount of calculation does not easily increase, and huge amounts of data can be processed in a short time.

[0108] Next, the text analysis data AD REFFor this, related words corresponding to the keyword are extracted (step S57). The extraction of related words corresponding to the keyword is based on the degree of similarity or proximity between the distributed representation vector of the keyword and the distributed representation vector of the word, and related words corresponding to the keyword are extracted. Then, by arranging the related words in descending order of similarity or ascending order of proximity, related word data is generated. Specifically, it is preferable to extract 1 or more and 10 or fewer related words for one keyword, and more preferably 2 or more and 5 or fewer related words. The related words may be, for example, words with a similarity of a predetermined value or more, words with a distance of a predetermined value or less, the top predetermined number of words with a high similarity, or the top predetermined number of words with a close distance. Depending on the keyword, the number of synonyms, antonyms, hypernyms, hyponyms, etc. is different. Therefore, depending on the keyword, the number of related words may be different. The text analysis data AD REF By extracting related words of the keyword from the words included in the text analysis data AD REF even when the text analysis data AD expresses the keyword in a unique notation, the notation can be extracted as a related word. Therefore, it is preferable because it can reduce search omissions due to notation fluctuations.

[0109] The similarity between two vectors can be obtained using cosine similarity, covariance, unbiased covariance, Pearson's product-moment correlation coefficient, etc. In particular, it is preferable to use cosine similarity. The distance between two vectors can be obtained using Euclidean distance, standard (standardized, average) Euclidean distance, Mahalanobis distance, Manhattan distance, Chebyshev distance, Minkowski distance, etc.

[0110] Next, weight data is assigned to the related words (step S58). The weight data assigned to each related word corresponds to the degree of relevance (similarity) between the keyword and the related word. Therefore, the weight data assigned to the related word is a value indicating the high degree of similarity or the short distance, or a value obtained by normalizing these. The weight data assigned to the related word is used for calculating the weight of the related word, which is later used when assigning scores to the search results. Specifically, the product of the normalized IDF of the keyword and the weight data of the related word corresponds to the weight of the related word. Note that the calculation of the weight of the related word only needs to be a value corresponding to the product, and a value corresponding to the intercept of the product may be added to the calculated weight value.

[0111] Using the above-described IDF data ID and vector data VD, a thesaurus dictionary data composed of a plurality of keywords and related words with weight data attached is generated (step S59). The generated thesaurus dictionary data can be stored in the database 50.

[0112] According to the flow of FIG. 4, thesaurus dictionary data can be generated based on a plurality of document data stored in the database. For each related word related to the characteristic word (keyword) in the document data, the similarity (weight) for each related word can be estimated by estimating with a numerical value normalized by the IDF data ID and the vector data VD. Therefore, a document search method capable of searching documents with high accuracy can be provided. Alternatively, with a simple input method, high-precision document search, particularly document search related to intellectual property, can be realized.

[0113] FIG. 5 is a diagram showing a flow for extracting a search keyword shown in step S15 described in FIG. 2.

[0114] First, document data (hereinafter referred to as document data TD) is input to the processing unit 30 via the input unit 20 (step S31). Step S31 corresponds to step S14 described above.

[0115] Next, perform word segmentation processing on the sentence data TD (step S32). After that, it is preferable to perform a process for correcting unnecessary word segmentation processing.

[0116] Next, perform morphological analysis on the sentence data TD that has undergone word segmentation processing (step S33).

[0117] Next, generate document analysis data (hereinafter referred to as document analysis data AD) for the data on which morphological analysis has been performed (step S34). In morphological analysis, a sentence written in natural language can be divided into morphemes (the smallest units that have meaning as a language), and the part-of-speech of the morphemes can be determined. As a result, for example, document analysis data AD can be obtained by extracting only nouns from the sentence data TD that has undergone word segmentation processing.

[0118] Next, refer to the IDF data calculated when generating the weight dictionary data or the thesaurus data, and acquire the IDF data ID corresponding to the words included in the document analysis data AD (step S35). By acquiring the IDF data ID normalized for each word, it is possible to extract search words, which are characteristic words for searching for a desired document, and the normalized IDF.

[0119] Next, extract search words based on the IDF (step S36). Words with a high IDF are characteristic words that rarely appear in the document analysis data AD.

[0120] According to the flow of FIG. 5, search words can be extracted based on the input sentence data. By estimating the characteristic words in the sentence data with numerical values normalized by the IDF, the characteristic words can be extracted as search words. Therefore, it is possible to provide a document search method that can search for documents with high accuracy. Alternatively, with a simple input method, it is possible to realize highly accurate document search, particularly for documents related to intellectual property.

[0121] FIG. 6A is a diagram schematically showing data of search words (SW) extracted from the text data TD described above. Table data 21TB schematically shows data of search words (SW). As the extracted search words, "Word A", "Word B", and "Word C" are exemplified.

[0122] FIG. 6B is a diagram schematically showing weighted dictionary data with weight data based on normalized IDF attached to each keyword (KW) generated from the plurality of document data described above. Table data 52TB schematically shows the weighted dictionary data. As the keywords, "Word A", "Word B", and "Word C" are exemplified, and the weight data of each keyword are "0.9", "0.9", and "0.8", respectively.

[0123] FIG. 6C is a diagram schematically showing synonym dictionary data in which related words are extracted for each keyword (KW) extracted from the plurality of document data described above, and weight data corresponding to similarity is attached to each related word (RW). Table data 53TB schematically shows the synonym dictionary data.

[0124] In Table 53TB, as keywords KW, "Word A", "Word B", "Word C", "Word D", and "Word E" are exemplified. As related words of "Word A", "Word X", "Word Y", "Word Z", and "Word a" are exemplified, and as the weight data of each related word, "0.9", "0.8", "0.6", and "0.5" are used. Similarly, as related words of "Word B", "Word b", "Word c", "Word d", and "Word e" are exemplified, and as the weight data of each related word, "0.5", "0.5", "0.45", and "0.3" are used. As related words of "Word C", "Word f", "Word g", "Word h", and "Word i" are exemplified, and as the weight data of each related word, "0.75", "0.75", "0.75", and "0.75" are used. As related words of "Word D", "Word j", "Word k", "Word m", and "Word n" are exemplified, and as the weight data of each related word, "0.5", "0.3", "0.3", and "0.1" are used. As related words of "Word E", "Word p", "Word q", "Word r", and "Word s" are exemplified, and as the weight data of each related word, "0.75", "0.65", "0.65", and "0.6" are used.

[0125] FIG. 7 is a diagram schematically showing search data created by referring to weight dictionary data and similar word dictionary data. In table data 61TB, the weights of "Word A", "Word B", and "Word C" shown in table data 21TB having search word SW are set to "0.9", "0.9", and "0.8" by referring to table data 52TB. Also, as related words corresponding to keyword KW, by referring to table data 53TB, for "Word A", "Word X", "Word Y", "Word Z", and "Word a" are exemplified, and their respective weights as related words are set to "0.9", "0.8", "0.6", and "0.5". Similarly, for "Word B", "Word b", "Word c", "Word d", and "Word e" are exemplified, and their respective weights as related words are set to "0.5", "0.5", "0.45", and "0.3". For "Word C", "Word f", "Word g", "Word h", and "Word i" are exemplified, and their respective weights as related words are set to "0.75", "0.75", "0.75", and "0.75".

[0126] The table data 61TB illustrated in FIG. 7 is displayed on a display device provided outside the document search system 10. The user can view the search data displayed on the display device provided outside the document search system 10 as illustrated in the table data 61TB in FIG. 7, and correct the weight data of words that are clearly not appropriate as related words or the weight data of related words with clearly high relevance.

[0127] For example, as illustrated in FIG. 8, in the table data 61TB illustrated in FIG. 7, in the case of "Word A", if the relevance of "Word a" is large according to the user's judgment, the weight of the related word is corrected from "0.5" to "1.0". Similarly, in the case of "Word B", if the relevance of "Word c" is small according to the user's judgment, the weight of the related word is corrected from "0.5" to "0.0". Similarly, in the case of "Word C", if the relevance of "Word h" is large according to the user's judgment, the weight of the related word is corrected from "0.75" to "1.0". Note that the related words with corrected weight data are hatched.

[0128] When the user makes the modification shown in FIG. 8, the search data (first search data: corresponding to table data 61TB) becomes modified search data (second search data: corresponding to table data 62TB).

[0129] Note that the update of the thesaurus dictionary data is not limited to the example shown in FIG. 8. For example, when modifying the weight data of related words from "0.5" to "1.0", it may be a modification considering the contribution rate. For example, a configuration may be adopted in which a value obtained by multiplying the contribution rate by the difference between the weight data before modification and the weight data after modification is added to the weight data before modification to obtain the weight data after modification. In the case of this configuration, assuming that the contribution rate is 0.1, the weight data before modification is 0.5, and the weight data after modification is 1.0, the weight data after modification is updated to 0.55 as "0.5 + 0.1×(1.0 - 0.5)". Therefore, when updating the thesaurus dictionary data, regardless of the modification content of a single user, it is possible to perform an update in response to the modifications of multiple users.

[0130] Also, FIG. 9 is a diagram schematically showing the thesaurus dictionary data updated when the search data shown in FIG. 8 is modified. The related word RW (hatched part) with modified weight data and the corresponding keyword KW shown in FIG. 8 are such that the thesaurus dictionary data is modified based on the weight data to be modified. Specifically, the table data 53TB schematically representing the thesaurus dictionary data before update shown in FIG. 9 can be updated as shown in table data 53TB_re.

[0131] As shown in FIG. 9, for the related word RW with updated weight data, the ranking of the related words linked to the keyword fluctuates. By updating the thesaurus dictionary data in this way, it is possible to provide a document search method capable of searching for documents taking into account the user's judgment criteria. Also, it is possible to provide a document search method capable of searching for documents with high accuracy. Alternatively, with a simple input method, it is possible to realize highly accurate document search, particularly for documents related to intellectual property.

[0132] FIG. 10 is a diagram showing a flow for explaining the update of the thesaurus dictionary data shown in step S22 described in FIG. 2.

[0133] First, the search data corrected by the user is stored in the storage unit via the input unit (step S61). Step S61 corresponds to step S20 described in FIG. 2 above.

[0134] Next, it is determined whether to perform periodic update of the thesaurus dictionary data (step S62). The periodic update is performed using a timer or the like. In the case of the update timing, the thesaurus dictionary data is updated (step S63). If not updated, the process ends. The update of the thesaurus dictionary data in step S63 is performed regardless of whether the search data is stored in step S61.

[0135] FIG. 11 is a diagram showing a flow for explaining the search execution shown in step S19 described in FIG. 2.

[0136] First, search data is created based on the search word (step S71). Step S71 corresponds to step S16 described above.

[0137] Next, the created search data is corrected (step S72). Step S72 corresponds to step S18 described above. In this way, by the user editing (correcting) the weight data, the search accuracy can be improved.

[0138] Next, scoring is performed on the reference text analysis data AD ref based on the weight data attached to the search data (step S73). The scoring process for a plurality of reference text analysis data AD ref will be described with reference to FIG. 12 and the like described later.

[0139] Next, the reference text analysis data AD refCreate ranking data based on the scores assigned to each of them (step S74).

[0140] The ranking data can include the rank, information (such as name and identification number) (Doc) of the reference text data TD ref and the score (Score), etc. If the reference text data TD ref is stored in the database 50 or the like, the ranking data preferably includes the file path to the reference text data TD ref Thus, the user can easily access the target document from the ranking data.

[0141] The higher the score of the reference text analysis data AD ref , the more it can be said that the text analysis data AD ref is related or similar to the text data TD.

[0142] The document search system according to one aspect of the present invention has a function of extracting search words based on text data and extracting keywords and related words of the keywords by referring to a thesaurus data and a weight dictionary data. Therefore, the user of the document search system according to one aspect of the present invention does not have to select the keywords used for the search by himself / herself. The user can directly input text data (text data) having a larger amount than the keywords into the document search system. Also, when the user himself / herself wants to select keywords and related words, there is no need to select them from scratch, and the user can refer to the keywords and related words extracted by the document search system and perform addition, modification, deletion, etc. of the keywords and related words. Therefore, the burden on the user in document search can be reduced, and the difference in search results due to the user's skill can be made less likely to occur.

[0143] FIG. 12 is a diagram showing a flow for explaining the scoring of the reference text analysis data AD ref based on the weight data attached to the search data shown in step S73 described in FIG. 11.

[0144] Unscored reference text analysis data AD ref Select one (step S81).

[0145] Next, in the reference text analysis data AD ref Determine whether the keyword KW hits (step S82). If it hits, proceed to step S85. If it does not hit, proceed to step S83.

[0146] Next, in the reference text analysis data AD ref Determine whether the related word RW corresponding to the keyword KW hits (step S83). If it hits, proceed to step S85. If it does not hit, proceed to step S84.

[0147] Next, determine whether all related words RW corresponding to the keyword KW have been searched (step S84). If they have been searched, proceed to step S86. If they have not been searched, proceed to step S83. For example, if there are two related words RW for the keyword KW and it was determined in the previous step S83 whether the first related word RW hits, return to step S83 to determine whether the second related word RW hits.

[0148] In step S85, add the weight corresponding to the hit word to the score. If it hits in step S82, add the weight data of the keyword KW to the score. If it hits in step S83, add the product of the weight data of the keyword KW x and the weight data of the related word RW to the score.

[0149] Next, determine whether all keywords KW have been searched (step S86). If they have been searched, proceed to step S87. If they have not been searched, proceed to step S82. For example, if there are two keywords KW and it was determined in the previous step S82 whether the first keyword KW hits, return to step S82 to determine whether the second keyword KW hits.

[0150] Next, all reference text analysis data ADref Determine whether scoring has been done for it (step S87). If all scoring is completed, the process ends. If not, proceed to step S81.

[0151] As described above, the search can be performed using the document search system 10.

[0152] As described above, in the document search system of the present embodiment, among the pre-prepared documents as the search targets, documents related or similar to the input document can be searched. Since it is not necessary for the user to select the keywords used for the search and the search can be performed using the text data with a larger volume than the keywords, the individual differences in search accuracy can be reduced, and the documents can be searched simply and with high accuracy. Further, in the document search system of the present embodiment, since the related words of the keywords are extracted from the pre-prepared documents, the unique notations included in the documents can also be extracted as related words, and the search omission can be reduced. Further, since the document search system of the present embodiment can rank and output the search results according to the degree of relevance or similarity, it is easy for the user to find the necessary documents from the search results and difficult to overlook them.

[0153] The present embodiment can be appropriately combined with other embodiments. Further, in this specification, when a plurality of configuration examples are shown in one embodiment, the configuration examples can be appropriately combined.

[0154] (Embodiment 2) In the present embodiment, a configuration example of a semiconductor device that can be used for a neural network will be described.

[0155] The semiconductor device of the present embodiment can be used, for example, in the processing unit of the document search system according to one aspect of the present invention.

[0156] As shown in FIG. 13A, the neural network NN can be composed of an input layer IL, an output layer OL, and an intermediate layer (hidden layer) HL. The input layer IL, the output layer OL, and the intermediate layer HL each have one or more neurons (units). Note that the intermediate layer HL may be one layer or two or more layers. A neural network having two or more intermediate layers HL can also be called a DNN (Deep Neural Network), and learning using a deep neural network can also be called deep learning.

[0157] Input data is input to each neuron in the input layer IL, the output signals of neurons in the previous layer or the next layer are input to each neuron in the intermediate layer HL, and the output signals of neurons in the previous layer are input to each neuron in the output layer OL. Note that each neuron may be connected to all neurons in the previous and subsequent layers (fully connected), or may be connected to some neurons.

[0158] FIG. 13B shows an example of an operation by a neuron. Here, a neuron N and two neurons in the previous layer that output signals to the neuron N are shown. The output x1 of the neuron in the previous layer and the output x2 of the neuron in the previous layer are input to the neuron N. Then, in the neuron N, after calculating the sum x1w1 + x2w2 of the multiplication result (x1w1) of the output x1 and the weight w1 and the multiplication result (x2w2) of the output x2 and the weight w2, a bias b is added as necessary to obtain a value a = x1w1 + x2w2 + b. Then, the value a is converted by the activation function h, and an output signal y = h(a) is output from the neuron N.

[0159] Thus, the operation by neurons includes an operation of adding the products of the outputs of neurons in the previous layer and the weights, that is, a product-sum operation (x1w1 + x2w2 as described above). This product-sum operation may be performed on software using a program, or may be performed by hardware. When performing the product-sum operation by hardware, a product-sum operation circuit can be used. As this product-sum operation circuit, a digital circuit or an analog circuit may be used. When using an analog circuit for the product-sum operation circuit, it is possible to reduce the circuit scale of the product-sum operation circuit, or improve the processing speed and reduce the power consumption by reducing the number of accesses to the memory.

[0160] The product-sum operation circuit may be constituted by transistors (also referred to as "Si transistors") including silicon (such as single-crystalline silicon) in the channel formation region, or may be constituted by transistors (also referred to as "OS transistors") including an oxide semiconductor which is a kind of metal oxide in the channel formation region. In particular, since the OS transistor has an extremely small off-current, it is suitable as a transistor constituting the memory of the product-sum operation circuit. Note that the product-sum operation circuit may be constituted using both Si transistors and OS transistors. Hereinafter, a configuration example of a semiconductor device having the function of the product-sum operation circuit will be described.

[0161] <Configuration Example of Semiconductor Device> FIG. 14 shows a configuration example of a semiconductor device MAC having a function of performing operations of a neural network. The semiconductor device MAC has a function of performing a product-sum operation on first data corresponding to the coupling strength (weight) between neurons and second data corresponding to input data. Note that the first data and the second data can each be analog data or multi-valued digital data (discrete data). Further, the semiconductor device MAC has a function of converting the data obtained by the product-sum operation by an activation function.

[0162] The semiconductor device MAC includes a cell array CA, a current source circuit CS, a current mirror circuit CM, a circuit WDD, a circuit WLD, a circuit CLD, an offset circuit OFST, and an activation function circuit ACTV.

[0163] The cell array CA includes a plurality of memory cells MC and a plurality of memory cells MCref. FIG. 14 shows a configuration example in which the cell array CA has memory cells MC (MC[1,1] to MC[m,n]) arranged in m rows and n columns (m and n are integers of 1 or more) and m memory cells MCref (MCref[1] to MCref[m]). The memory cell MC has a function of storing first data. The memory cell MCref has a function of storing reference data used for the sum-of-products operation. Note that the reference data can be analog data or multi-valued digital data.

[0164] The memory cell MC[i,j] (where i is an integer from 1 to m and j is an integer from 1 to n) is connected to a wiring WL[i], a wiring RW[i], a wiring WD[j], and a wiring BL[j]. The memory cell MCref[i] is connected to a wiring WL[i], a wiring RW[i], a wiring WDref, and a wiring BLref. Here, the current flowing between the memory cell MC[i,j] and the wiring BL[j] is denoted as I MC[i,j] and the current flowing between the memory cell MCref[i] and the wiring BLref is denoted as I MCref[i] as described above.

[0165] Specific configuration examples of the memory cell MC and the memory cell MCref are shown in FIG. 15. FIG. 15 shows, as representative examples, the memory cells MC[1,1], MC[2,1] and the memory cells MCref[1], MCref[2], but the same configuration can be used for other memory cells MC and memory cells MCref. Each of the memory cell MC and the memory cell MCref has a transistor Tr11, a transistor Tr12, and a capacitor element C11. Here, a case where the transistor Tr11 and the transistor Tr12 are n-channel type transistors will be described.

[0166] In the memory cell MC, the gate of the transistor Tr11 is connected to the wiring WL, one of the source or drain is connected to the gate of the transistor Tr12 and the first electrode of the capacitor element C11, and the other of the source or drain is connected to the wiring WD. One of the source or drain of the transistor Tr12 is connected to the wiring BL, and the other of the source or drain is connected to the wiring VR. The second electrode of the capacitor element C11 is connected to the wiring RW. The wiring VR is a wiring having a function of supplying a predetermined potential. Here, as an example, the case where a low power supply potential (such as a ground potential) is supplied from the wiring VR will be described.

[0167] A node connected to one of the source or drain of the transistor Tr11, the gate of the transistor Tr12, and the first electrode of the capacitor element C11 is defined as the node NM. Also, the nodes NM of the memory cells MC[1,1] and MC[2,1] are denoted as nodes NM[1,1] and NM[2,1], respectively.

[0168] The memory cell MCref also has the same configuration as the memory cell MC. However, the memory cell MCref is connected to the wiring WDref instead of the wiring WD, and is connected to the wiring BLref instead of the wiring BL. Also, in the memory cells MCref[1] and MCref[2], a node connected to one of the source or drain of the transistor Tr11, the gate of the transistor Tr12, and the first electrode of the capacitor element C11 is denoted as the node NMref[1] and NMref[2], respectively.

[0169] The node NM and the node NMref each function as a holding node of the memory cell MC and the memory cell MCref. The first data is held in the node NM, and the reference data is held in the node NMref. Also, currents I MC[1,1] , I MC[2,1] flow into the transistors Tr12 of the memory cells MC[1,1] and MC[2,1] from the wiring BL, respectively. Also, currents I MCref[1] , IMCref[2] flows.

[0170] Since the transistor Tr11 has a function of holding the potential of the node NM or the node NMref, it is preferable that the off-current of the transistor Tr11 is small. Therefore, it is preferable to use an OS transistor having an extremely small off-current as the transistor Tr11. Thereby, fluctuations in the potential of the node NM or the node NMref can be suppressed, and the calculation accuracy can be improved. In addition, the frequency of the operation for refreshing the potential of the node NM or the node NMref can be kept low, and the power consumption can be reduced.

[0171] The transistor Tr12 is not particularly limited, and for example, an Si transistor or an OS transistor can be used. When an OS transistor is used for the transistor Tr12, the transistor Tr12 can be manufactured using the same manufacturing apparatus as the transistor Tr11, and the manufacturing cost can be suppressed. Note that the transistor Tr12 may be an n-channel type or a p-channel type.

[0172] The current source circuit CS is connected to the wirings BL[1] to BL[n] and the wiring BLref. The current source circuit CS has a function of supplying current to the wirings BL[1] to BL[n] and the wiring BLref. Note that the current value supplied to the wirings BL[1] to BL[n] may be different from the current value supplied to the wiring BLref. Here, the current supplied from the current source circuit CS to the wirings BL[1] to BL[n] is I C , and the current supplied from the current source circuit CS to the wiring BLref is I Cref is denoted.

[0173] The current mirror circuit CM has wirings IL[1] to IL[n] and wiring ILref. The wirings IL[1] to IL[n] are each connected to the wirings BL[1] to BL[n], and the wiring ILref is connected to the wiring BLref. Here, the connection points of the wirings IL[1] to IL[n] and the wirings BL[1] to BL[n] are denoted as nodes NP[1] to NP[n]. Also, the connection point of the wiring ILref and the wiring BLref is denoted as node NPref.

[0174] The current mirror circuit CM has a function of flowing a current I CM corresponding to the potential at node NPref through the wiring ILref, and a function of flowing this current I CM also through the wirings IL[1] to IL[n]. FIG. 14 shows an example in which the current I CM is discharged from the wiring BLref to the wiring ILref, and the current I CM is discharged from the wirings BL[1] to BL[n] to the wirings IL[1] to IL[n]. Also, the currents flowing from the current mirror circuit CM to the cell array CA via the wirings BL[1] to BL[n] are denoted as I B [1] to I B [n]. Also, the current flowing from the current mirror circuit CM to the cell array CA via the wiring BLref is denoted as I Bref .

[0175] The circuit WDD is connected to the wirings WD[1] to WD[n] and the wiring WDref. The circuit WDD has a function of supplying a potential corresponding to the first data stored in the memory cell MC to the wirings WD[1] to WD[n]. Also, the circuit WDD has a function of supplying a potential corresponding to the reference data stored in the memory cell MCref to the wiring WDref. The circuit WLD is connected to the wirings WL[1] to WL[m]. The circuit WLD has a function of supplying a signal for selecting the memory cell MC or the memory cell MCref that writes data to the wirings WL[1] to WL[m]. The circuit CLD is connected to the wirings RW[1] to RW[m]. The circuit CLD has a function of supplying a potential corresponding to the second data to the wirings RW[1] to RW[m].

[0176] The offset circuit OFST is connected to wirings BL[1] to BL[n] and wirings OL[1] to OL[n]. The offset circuit OFST has a function of detecting the amount of current flowing from the wirings BL[1] to BL[n] into the offset circuit OFST and / or the amount of change in the current flowing from the wirings BL[1] to BL[n] into the offset circuit OFST. Further, the offset circuit OFST has a function of outputting the detection result to the wirings OL[1] to OL[n]. Note that the offset circuit OFST may output a current corresponding to the detection result to the wiring OL, or may convert the current corresponding to the detection result into a voltage and output the voltage to the wiring OL. The current flowing between the cell array CA and the offset circuit OFST is represented as I α [1] to I α [n].

[0177] A configuration example of the offset circuit OFST is shown in FIG. 16. The offset circuit OFST shown in FIG. 16 has circuits OC[1] to OC[n]. Further, each of the circuits OC[1] to OC[n] has a transistor Tr21, a transistor Tr22, a transistor Tr23, a capacitive element C21, and a resistive element R1. The connection relationship of each element is as shown in FIG. 16. Note that a node connected to the first electrode of the capacitive element C21 and the first terminal of the resistive element R1 is defined as node Na. Also, a node connected to the second electrode of the capacitive element C21, either the source or the drain of the transistor Tr21, and the gate of the transistor Tr22 is defined as node Nb.

[0178] The wiring VrefL has the function of supplying the potential Vref, the wiring VaL has the function of supplying the potential Va, and the wiring VbL has the function of supplying the potential Vb. Also, the wiring VDDL has the function of supplying the potential VDD, and the wiring VSSL has the function of supplying the potential VSS. Here, the case where the potential VDD is the high power supply potential and the potential VSS is the low power supply potential will be described. Also, the wiring RST has the function of supplying a potential for controlling the conduction state of the transistor Tr21. A source follower circuit is configured by the transistor Tr22, the transistor Tr23, the wiring VDDL, the wiring VSSL, and the wiring VbL.

[0179] Next, the operation examples of the circuits OC[1] to OC[n] will be described. Here, as a representative example, the operation example of the circuit OC[1] will be described, but the circuits OC[2] to OC[n] can also be operated in the same manner. First, when a first current flows through the wiring BL[1], the potential of the node Na becomes a potential corresponding to the first current and the resistance value of the resistance element R1. At this time, the transistor Tr21 is in the on state, and the potential Va is supplied to the node Nb. Then, the transistor Tr21 becomes the off state.

[0180] Next, when a second current flows through the wiring BL[1], the potential of the node Na changes to a potential corresponding to the second current and the resistance value of the resistance element R1. At this time, the transistor Tr21 is in the off state, and since the node Nb is in the floating state, the potential of the node Nb changes by capacitive coupling as the potential of the node Na changes. Here, let the change in the potential of the node Na be ΔV Na and assuming the capacitive coupling coefficient is 1, the potential of the node Nb becomes Va + ΔV Na Then, assuming the threshold voltage of the transistor Tr22 is V th a potential Va + ΔV Na - V th is output from the wiring OL[1]. Here, by setting Va = V th a potential ΔV Na can be output from the wiring OL[1].

[0181] The potential ΔV Nais determined according to the change amount from the first current to the second current, the resistance value of the resistance element R1, and the potential Vref. Here, since the resistance value of the resistance element R1 and the potential Vref are known, the potential ΔV Na can be used to obtain the change amount of the current flowing from the wiring BL.

[0182] The amount of current detected by the offset circuit OFST as described above and / or the signal corresponding to the change amount of the current are input to the activation function circuit ACTV via the wirings OL[1] to OL[n].

[0183] The activation function circuit ACTV is connected to the wirings OL[1] to OL[n] and the wirings NIL[1] to NIL[n]. The activation function circuit ACTV has a function of performing an operation for converting the signal input from the offset circuit OFST according to a predefined activation function. As the activation function, for example, a sigmoid function, a tanh function, a softmax function, a ReLU function, a threshold function, etc. can be used. The signal converted by the activation function circuit ACTV is output to the wirings NIL[1] to NIL[n] as output data.

[0184] <Operation example of the semiconductor device> Using the semiconductor device MAC described above, a multiplication and addition operation of the first data and the second data can be performed. Hereinafter, an operation example of the semiconductor device MAC when performing a multiplication and addition operation will be described.

[0185] FIG. 17 shows a timing chart of an operation example of the semiconductor device MAC. FIG. 17 shows the potential transitions of the wirings WL[1], WL[2], WD[1], WDref, nodes NM[1,1], NM[2,1], NMref[1], NMref[2], wiring RW[1], and wiring RW[2] in FIG. 15, and the current I B [1] - I α [1], and the current I Bref The value transitions of are shown. The current I B [1] - I α[1] corresponds to the sum of the currents flowing from wiring BL[1] to memory cells MC[1,1] and MC[2,1].

[0186] Here, as a representative example, the operation will be described by focusing on the memory cells MC[1,1] and MC[2,1] and the memory cells MCref[1] and MCref[2] shown in FIG. 15. However, other memory cells MC and memory cells MCref can be operated in the same manner.

[0187] [Storage of First Data] First, in the period from time T01 to time T02, the potential of wiring WL[1] becomes high level, the potential of wiring WD[1] becomes higher than the ground potential (GND) by V PR -V W[1,1] and the potential of wiring WDref becomes higher than the ground potential by V PR . Also, the potentials of wiring RW[1] and wiring RW[2] become the reference potential (REFP). Here, the potential V W[1,1] corresponds to the potential of the first data stored in memory cell MC[1,1]. Also, the potential V PR corresponds to the potential of the reference data. As a result, the transistor Tr11 included in memory cell MC[1,1] and memory cell MCref[1] turns on, and the potential of node NM[1,1] becomes V PR -V W[1,1] , and the potential of node NMref[1] becomes V PR .

[0188] At this time, the current I MC[1,1],0 flowing from wiring BL[1] to transistor Tr12 of memory cell MC[1,1] can be expressed by the following equation. Here, k is a constant determined by the channel length, channel width, mobility, and capacitance of the gate insulating film of transistor Tr12. Also, V th is the threshold voltage of transistor Tr12.

[0189] I MC[1,1],0 =k(V PR -V W[1,1] -V th ) 2(E1)

[0190] Also, the current I flowing from the wiring BLref to the transistor Tr12 of the memory cell MCref[1] MCref[1],0 can be expressed by the following equation.

[0191] I MCref[1],0 =k(V PR -V th ) 2 (E2)

[0192] Next, during the period from time T02 to time T03, the potential of the wiring WL[1] becomes low level. As a result, the transistor Tr11 included in the memory cell MC[1,1] and the memory cell MCref[1] turns off, and the potentials of the nodes NM[1,1] and NMref[1] are held.

[0193] As described above, it is preferable to use an OS transistor as the transistor Tr11. Thereby, the leakage current of the transistor Tr11 can be suppressed, and the potentials of the nodes NM[1,1] and NMref[1] can be accurately held.

[0194] Next, during the period from time T03 to time T04, the potential of the wiring WL[2] becomes high level, the potential of the wiring WD[1] becomes higher than the ground potential by V PR -V W[2,1] and the potential of the wiring WDref becomes higher than the ground potential by V PR . Note that the potential V W[2,1] corresponds to the potential of the first data stored in the memory cell MC[2,1]. As a result, the transistors Tr11 included in the memory cell MC[2,1] and the memory cell MCref[2] turn on, the potential of the node NM[2,1] becomes V PR -V W[2,1] and the potential of the node NMref[2] becomes V PR .

[0195] At this time, the current I flowing from the wiring BL[1] to the transistor Tr12 of the memory cell MC[2,1]MC[2,1],0 can be expressed by the following formula.

[0196] I MC[2,1],0 = k(V PR - V W[2,1] - V th )(E3) 2 (E3)

[0197] Also, the current I flowing from the wiring BLref to the transistor Tr12 of the memory cell MCref[2] MCref[2],0 can be expressed by the following formula.

[0198] I MCref[2],0 = k(V PR - V th )(E4) 2 (E4)

[0199] Next, during the period from time T04 to time T05, the potential of the wiring WL[2] becomes low level. As a result, the transistors Tr11 of the memory cells MC[2,1] and MCref[2] are turned off, and the potentials of the nodes NM[2,1] and NMref[2] are held.

[0200] Through the above operations, the first data is stored in the memory cells MC[1,1] and MC[2,1], and the reference data is stored in the memory cells MCref[1] and MCref[2].

[0201] Here, consider the currents flowing through the wiring BL[1] and the wiring BLref during the period from time T04 to time T05. A current is supplied from the current source circuit CS to the wiring BLref. Also, the current flowing through the wiring BLref is discharged to the current mirror circuit CM and the memory cells MCref[1] and MCref[2]. Let the current supplied from the current source circuit CS to the wiring BLref be I Cref , and the current discharged from the wiring BLref to the current mirror circuit CM be I CM,0 . Then, the following equation holds.

[0202] I Cref - I CM,0 = I MCref[1],0+I MCref[2],0 (E5)

[0203] The wiring BL[1] is supplied with current from the current source circuit CS. Also, the current flowing through the wiring BL[1] is discharged to the current mirror circuit CM, the memory cells MC[1,1], and MC[2,1]. Also, current flows from the wiring BL[1] to the offset circuit OFST. Let the current supplied from the current source circuit CS to the wiring BL[1] be I C,0 , and the current flowing from the wiring BL[1] to the offset circuit OFST be I α,0 . Then, the following equation holds.

[0204] I C - I CM,0 = I MC[1,1],0 + I MC[2,1],0 + I α,0 (E6)

[0205] [Sum-of-products operation of the first data and the second data] Next, in the period from time T05 to time T06, the potential of the wiring RW[1] becomes a potential higher than the reference potential by V X[1] . At this time, the potential V X[1] is supplied to the respective capacitive elements C11 of the memory cell MC[1,1] and the memory cell MCref[1], and the potential of the gate of the transistor Tr12 rises due to capacitive coupling. Note that the potential V X[1] is the potential corresponding to the second data supplied to the memory cell MC[1,1] and the memory cell MCref[1].

[0206] The amount of change in the potential of the gate of the transistor Tr12 is a value obtained by multiplying the amount of change in the potential of the wiring RW by the capacitive coupling coefficient determined by the configuration of the memory cell. The capacitive coupling coefficient is calculated from the capacitance of the capacitive element C11, the gate capacitance of the transistor Tr12, and parasitic capacitance, etc. Hereinafter, for convenience, it will be described assuming that the amount of change in the potential of the wiring RW and the amount of change in the potential of the gate of the transistor Tr12 are the same, that is, the capacitive coupling coefficient is 1. Actually, the potential V X may be determined in consideration of the capacitive coupling coefficient.

[0207] When a potential V is supplied to the capacitance element C11 of the memory cell MC[1,1] and the memory cell MCref[1], the potentials of the nodes NM[1,1] and NMref[1] each become V X[1] and rise X[1]

[0208] Here, during the period from time T05 to time T06, the current I flowing from the wiring BL[1] to the transistor Tr12 of the memory cell MC[1,1] MC[1,1],1 can be expressed by the following equation

[0209] I MC[1,1],1 = k(V PR - V W[1,1] + V X[1] - V th )(E7) 2 (E7)

[0210] That is, by supplying the potential V X[1] to the wiring RW[1], the current flowing from the wiring BL[1] to the transistor Tr12 of the memory cell MC[1,1] is ΔI MC[1,1] = I MC[1,1],1 - I MC[1,1],0 and increases

[0211] Also, during the period from time T05 to time T06, the current I flowing from the wiring BLref to the transistor Tr12 of the memory cell MCref[1] MCref[1],1 can be expressed by the following equation

[0212] I MCref[1],1 = k(V PR + V X[1] - V th )(E8) 2 (E8)

[0213] That is, by supplying the potential V X[1] to the wiring RW[1], the current flowing from the wiring BLref to the transistor Tr12 of the memory cell MCref[1] is ΔI MCref[1] = I MCref[1],1 - I MCref[1],0 and increases

[0214] ​ Next, consider the currents flowing through wiring BL[1] and wiring BLref. A current I is supplied to wiring BLref from current source circuit CS. Also, the current flowing through wiring BLref is discharged to current mirror circuit CM, memory cells MCref[1] and MCref[2]. Let the current discharged from wiring BLref to current mirror circuit CM be I Cref . Then, the following equation holds. CM,1

[0215] I Cref - I CM,1 = I MCref[1],1 + I MCref[2],0 (E9)

[0216] A current I is supplied to wiring BL[1] from current source circuit CS. Also, the current flowing through wiring BL[1] is discharged to current mirror circuit CM, memory cells MC[1,1] and MC[2,1]. Further, a current also flows from wiring BL[1] to offset circuit OFST. Let the current flowing from wiring BL[1] to offset circuit OFST be I C . Then, the following equation holds. α,1

[0217] I C - I CM,1 = I MC[1,1],1 + I MC[2,1],1 + I α,1 (E10)

[0218] And from equations (E1) to (E10), the difference (differential current ΔI α,0 ) between current I α,1 and current I α can be expressed by the following equation.

[0219] ΔI α = I α,1 - I α,0 = 2kV W[1,1] V X[1] (E11)

[0220] Thus, differential current ΔI α is related to potentials V W[1,1] and V X[1] ​​It becomes a value corresponding to the product.

[0221] Thereafter, in the period from time T06 to time T07, the potential of the wiring RW[1] becomes the reference potential, and the potentials of the nodes NM[1,1] and NMref[1] become the same as those in the period from time T04 to time T05.

[0222] Next, in the period from time T07 to time T08, the potential of the wiring RW[1] becomes a potential higher than the reference potential by V X[1] and the potential of the wiring RW[2] becomes a potential higher than the reference potential by V X[2] As a result, the potential V X[1] is supplied to each of the capacitive elements C11 of the memory cell MC[1,1] and the memory cell MCref[1], and the potentials of the nodes NM[1,1] and NMref[1] rise to V X[1] respectively due to capacitive coupling. Also, the potential V X[2] is supplied to each of the capacitive elements C11 of the memory cell MC[2,1] and the memory cell MCref[2], and the potentials of the nodes NM[2,1] and NMref[2] rise to V X[2] respectively due to capacitive coupling.

[0223] Here, in the period from time T07 to time T08, the current I MC[2,1],1 flowing from the wiring BL[1] to the transistor Tr12 of the memory cell MC[2,1] can be expressed by the following equation.

[0224] I MC[2,1],1 =k(V PR -V W[2,1] +V X[2] -V th ) 2 (E12)

[0225] That is, by supplying the potential V X[2] to the wiring RW[2], the current flowing from the wiring BL[1] to the transistor Tr12 of the memory cell MC[2,1] increases by ΔI MC[2,1] =I MC[2,1],1 -I MC[2,1],0 .

[0226] Also, during the period from time T07 to time T08, the current I flowing from the wiring BLref to the transistor Tr12 of the memory cell MCref[2] MCref[2],1 can be expressed by the following equation.

[0227] I MCref[2],1 = k(V PR + V X[2] - V th )(E13) 2 (E13)

[0228] That is, by supplying the potential V X[2] to the wiring RW[2], the current flowing from the wiring BLref to the transistor Tr12 of the memory cell MCref[2] is ΔI MCref[2] = I MCref[2],1 - I MCref[2],0 and increases.

[0229] Also, consider the currents flowing through the wiring BL[1] and the wiring BLref. A current I Cref is supplied from the current source circuit CS to the wiring BLref. Also, the current flowing through the wiring BLref is discharged to the current mirror circuit CM, the memory cells MCref[1], MCref[2]. Let the current discharged from the wiring BLref to the current mirror circuit CM be I CM,2 . Then, the following equation holds.

[0230] I Cref - I CM,2 = I MCref[1],1 + I MCref[2],1 (E14)

[0231] A current I C is supplied from the current source circuit CS to the wiring BL[1]. Also, the current flowing through the wiring BL[1] is discharged to the current mirror circuit CM, the memory cells MC[1,1], MC[2,1]. Further, a current also flows from the wiring BL[1] to the offset circuit OFST. Let the current flowing from the wiring BL[1] to the offset circuit OFST be I α,2 . Then, the following equation holds.

[0232] I C-I CM,2 =I MC[1,1],1 +I MC[2,1],1 +I α,2 (E15)

[0233] Then, from Equation (E1) to Equation (E8), and Equation (E12) to Equation (E15), the current I α,0 and the current I α,2 the difference (differential current ΔI α ) can be expressed by the following equation.

[0234] ΔI α =I α,2 -I α,0 =2k(V W[1,1] V X[1] +V W[2,1] V X[2] ) (E16)

[0235] Thus, the differential current ΔI α is a value corresponding to the result of adding the product of the potential V W[1,1] and the potential V X[1] and the product of the potential V W[2,1] and the potential V X[2] .

[0236] Thereafter, during the period from time T08 to time T09, the potentials of the wirings RW[1], [2] become the reference potential, and the potentials of the nodes NM[1,1], NM[2,1] and the nodes NMref[1], NMref[2] are the same as those during the period from time T04 to time T05.

[0237] As shown in Equation (E11) and Equation (E16), the differential current ΔI α input to the offset circuit OFST can be calculated from an equation having a term of the product of the potential V W corresponding to the first data (weight) and the potential V X corresponding to the second data (input data). That is, by measuring the differential current ΔI α with the offset circuit OFST, the result of the sum-of-products operation of the first data and the second data can be obtained.

[0238] Note that in the above description, particular attention was paid to the memory cells MC[1,1], MC[2,1] and the reference memory cells MCref[1], MCref[2]. However, the number of memory cells MC and reference memory cells MCref can be arbitrarily set. When the number of rows m of the memory cells MC and the reference memory cells MCref is an arbitrary number i, the differential current ΔIα can be expressed by the following equation.

[0239] ΔI α =2kΣ i V W[i,1] V X[i] (E17)

[0240] Also, by increasing the number of columns n of the memory cells MC and the reference memory cells MCref, the number of sum-of-products operations executed in parallel can be increased.

[0241] As described above, by using the semiconductor device MAC, the sum-of-products operation between the first data and the second data can be performed. Note that by using the configuration shown in FIG. 15 for the memory cells MC and the reference memory cells MCref, a sum-of-products operation circuit can be configured with a small number of transistors. Therefore, the circuit scale of the semiconductor device MAC can be reduced.

[0242] When the semiconductor device MAC is used for operations in a neural network, the number of rows m of the memory cells MC can correspond to the number of input data supplied to one neuron, and the number of columns n of the memory cells MC can correspond to the number of neurons. For example, consider the case of performing a sum-of-products operation using the semiconductor device MAC in the intermediate layer HL shown in FIG. 13A. At this time, the number of rows m of the memory cells MC can be set to the number of input data (the number of neurons in the input layer IL) supplied from the input layer IL, and the number of columns n of the memory cells MC can be set to the number of neurons in the intermediate layer HL.

[0243] Note that the structure of the neural network to which the semiconductor device MAC is applied is not particularly limited. For example, the semiconductor device MAC can also be used in a convolutional neural network (CNN), a recurrent neural network (RNN), an autoencoder, a Boltzmann machine (including a restricted Boltzmann machine), etc.

[0244] As described above, by using the semiconductor device MAC, the multiplication-accumulation operation of the neural network can be performed. Furthermore, by using the memory cells MC and MCref shown in FIG. 15 in the cell array CA, an integrated circuit capable of improving the calculation accuracy, reducing the power consumption, or reducing the circuit scale can be provided.

[0245] This embodiment can be appropriately combined with other embodiments.

[0246] (Supplementary Note Regarding the Descriptions in this Specification, etc.) The above embodiments and the descriptions of each configuration in the embodiments are appended below.

[0247] The configurations shown in each embodiment can be appropriately combined with the configurations shown in other embodiments or examples to form an aspect of the present invention. Also, when a plurality of configuration examples are shown in one embodiment, the configuration examples can be appropriately combined.

[0248] Note that the content described in one embodiment (even part of the content) can be applied to, combined with, or replaced with the content described in another part (even part of the content) of the same embodiment and / or the content described in one or more other embodiments (even part of the content).

[0249] Note that the content described in the embodiments refers to the content described using various figures in each embodiment or the content described using the text described in the specification.

[0250] Note that the figures (which may be partial) described in one embodiment can be combined with other parts of the figure, other figures (which may be partial) described in that embodiment, and / or figures (which may be partial) described in one or more other embodiments to form even more figures.

[0251] Also, in this specification and the like, in a block diagram, components are classified by function and shown as independent blocks. However, in an actual circuit or the like, it is difficult to separate components by function, and there may be cases where a single circuit involves multiple functions or a single function involves multiple circuits. Therefore, the blocks in the block diagram are not limited to the components described in the specification and can be appropriately rephrased according to the situation.

[0252] Also, in the drawings, the size, layer thickness, or area is shown in an arbitrary size for convenience of explanation. Therefore, it is not necessarily limited to that scale. Note that the drawings are shown schematically for clarity and are not limited to the shapes or values shown in the drawings. For example, it is possible to include variations in signals, voltages, or currents due to noise, or variations in signals, voltages, or currents due to timing shifts.

[0253] Also, the positional relationship of the components illustrated in the drawings and the like is relative. Therefore, when explaining the components with reference to the drawings, terms such as "above" and "below" indicating the positional relationship may be used for convenience. The positional relationship of the components is not limited to the description in this specification and can be appropriately rephrased according to the situation.

[0254] In this specification and the like, when explaining the connection relationship of a transistor, the notation "one of the source or the drain" (or the first electrode, or the first terminal) and the other of the source and the drain are referred to as "the other of the source or the drain" (or the second electrode, or the second terminal). This is because the source and the drain of a transistor change depending on the structure of the transistor, the operating conditions, etc. Regarding the names of the source and the drain of a transistor, they can be appropriately rephrased according to the situation, such as the source (drain) terminal, the source (drain) electrode, etc.

[0255] Also, in this specification and the like, the terms "electrode" and "wiring" do not functionally limit these components. For example, an "electrode" may be used as part of a "wiring", and vice versa. Furthermore, the terms "electrode" and "wiring" also include cases where a plurality of "electrodes" and "wirings" are integrally formed.

[0256] Also, in this specification and the like, voltage and potential can be appropriately rephrased. Voltage is the potential difference from a reference potential. For example, if the reference potential is the ground voltage (earthing voltage), the voltage can be rephrased as potential. The ground potential does not necessarily mean 0V. Note that potential is relative, and depending on the reference potential, the potential applied to a wiring or the like may be changed.

[0257] Also, in this specification and the like, a node can be rephrased as a terminal, a wiring, an electrode, a conductive layer, a conductor, an impurity region, etc. according to the circuit configuration, the device structure, etc. Also, a terminal, a wiring, etc. can be rephrased as a node.

[0258] In this specification and the like, "A and B are connected" means that A and B are electrically connected. Here, "A and B are electrically connected" means a connection in which, when an object (such as a switch, a transistor element, a diode or other element, or a circuit including the element and wiring) exists between A and B, transmission of an electrical signal between A and B is possible. Note that when A and B are electrically connected, it includes the case where A and B are directly connected. Here, "A and B are directly connected" means a connection in which, without passing through the above object, transmission of an electrical signal between A and B is possible via wiring (or an electrode) or the like between A and B. In other words, direct connection means a connection that can be regarded as the same circuit diagram when represented by an equivalent circuit.

[0259] In this specification and the like, a switch refers to something that can be in a conductive state (on state) or a non-conductive state (off state) and has a function of controlling whether or not to allow current to flow. Or, a switch refers to something that has a function of selecting and switching a path through which current flows.

[0260] In this specification and the like, the channel length refers to, for example, in a top view of a transistor, the distance between the source and the drain in a region where a semiconductor (or the portion of the semiconductor where current flows when the transistor is in the on state) and the gate overlap, or in a region where a channel is formed.

[0261] In this specification and the like, the channel width refers to, for example, in a region where a semiconductor (or the portion of the semiconductor where current flows when the transistor is in the on state) and the gate electrode overlap, or in a region where a channel is formed, the length of the portion where the source and the drain face each other.

[0262] Note that in this specification and the like, terms such as "film" and "layer" can be interchanged with each other depending on the case or the situation. For example, in some cases, the term "conductive layer" can be changed to the term "conductive film". Or, for example, in some cases, the term "insulating film" can be changed to the term "insulating layer".

Description of Symbols

[0263] C11: Capacitance element, C21: Capacitance element, R1: Resistance element, Tr11: Transistor, Tr12: Transistor, Tr21: Transistor, Tr22: Transistor, Tr23: Transistor, 10: Document search system, 20: Input section, 21: Document data, 21TB: Table data, 30: Processing section, 40: Storage section, 50: Database, 51: Reference document data, 52: Weight dictionary data, 52TB: Table data, 53: Similar word search data, 53TB: Table data, 53TB_re: Table data, 60: Output section, 61: Search data, 61TB: Table data, 62: Search data, 62TB: Table data, 70: Transmission path

Claims

1. An input unit, a database, a storage unit, and a processing unit, wherein the database has a function of storing a plurality of reference document data, weight dictionary data, and synonym dictionary data, and the processing unit has a function of generating the weight dictionary data and the synonym dictionary data based on the reference document data, a function of generating document analysis data from the document data input to the input unit, a function of extracting a search word from among the words included in the document analysis data, and a function of generating first search data based on the search word and the weight dictionary data and the synonym dictionary data, wherein the storage unit has a function of storing second search data generated by the user modifying the first search data, and the processing unit has a function of updating the synonym dictionary data by adding a value obtained by multiplying the difference between the first search data and the second search data by the contribution rate per user to the first search data, and the processing unit has a function of generating reference document analysis data from the reference document data, and a function of extracting a plurality of keywords and related words corresponding to the keywords from among the words included in the reference document analysis data, wherein the weight dictionary data extracts the appearance frequency of the keyword from among the words included in the reference document analysis data, and is data generated by assigning a first weight corresponding to the appearance frequency to each of the keywords, wherein the first weight is a value based on the inverse document frequency of the keyword in the reference document analysis data, a document search system.

2. In claim 1, wherein the synonym dictionary data is data generated by assigning a second weight to each of the related words, a document search system.

3. In claim 2, wherein the second weight is the product of a value based on the similarity or distance between the distribution representation vector of the related word and the distribution representation vector of the keyword and the first weight of the keyword, a document search system.

4. In claim 3, wherein the distribution representation vector is a vector generated using a neural network, a document search system.

5. In any one of claims 1 to 4, wherein the processing unit has a transistor, and the transistor has a metal oxide in a channel formation region, a document search system.

6. In any one of claims 1 to 4, the processing unit has a transistor, the transistor has silicon in a channel formation region, a document retrieval system.

Citation Information

Patent Citations

  • Document retrieving method / device using syntax information of natural language

    JP1993342255A

  • Document search system

    JP2010003015A

  • Document retrieval system, information processing apparatus, and program

    JP2011090463A

  • Analyzer and analyzing method

    JP2015164008A

  • Information retrieval system, intellectual property information retrieval system, information retrieval method and intellectual property information retrieval method

    JP2018206376A