Contents classifying method
A computer-based content classification system using machine learning and a graphical user interface addresses inefficiencies in existing methods by enabling efficient and accurate classification of large volumes of content.
Patent Information
- Application Number
- JP2025140625
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-13
- Filing Date
- 2025-08-26
- Publication Date
- 2025-10-24
AI Technical Summary
Existing content classification methods face challenges in efficiently classifying large volumes of content due to reliance on user knowledge and experience, leading to variations in accuracy and efficiency, and require excessive user burden in generating training data.
A computer-based content classification system using machine learning to generate classification models, supported by a graphical user interface, allows for interactive classification and model generation, reducing user burden and improving accuracy.
The system enables efficient and accurate classification of content by leveraging machine learning, reducing user effort and minimizing variations in classification accuracy.
Smart Images

Figure 2025161942000001_ABST
Abstract
Description
[Technical Field]
[0001] One aspect of the present invention is a content classification method using a computer device, Classification systems, methods for generating classification models, and graphical user interfaces do.
[0002] One aspect of the present invention relates to a computer device. Digitalized content (text data, image data, audio data, In particular, one aspect of the present invention relates to a method for classifying a collection of content. This invention relates to a content classification system that efficiently classifies content using machine learning. One aspect of the present invention is a graphical user interface managed by a computer system through a program. Content classification method using interface, content classification system, and classification model This relates to a method for generating a del. [Background technology]
[0003] Users can easily find information on topics they specify from a collection of content. However, it is difficult to find the content that meets the desired criteria from a large amount of content. When classifying content, the content is classified according to the knowledge or experience of the individual. Similar results will vary.
[0004] In recent years, the results of content classification based on the knowledge and experience of individuals have been used as supervised data. It has been proposed to provide the data to a computer system and have it learn how to classify content by machine learning. For example, in Patent Document 1, documents highly related to topics specified by the user are searched. A machine learning approach to determining the target is disclosed. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-104630 Summary of the Invention [Problem to be solved by the invention]
[0006] A collection of content may be classified according to a purpose. This section explains the case where the content is a patent. Each patent is assigned a unique patent number. Therefore, in the following, the content may be referred to as the patent number. In addition, in the content classification method handled in one aspect of the present invention, multiple patent numbers are assigned to The focus is on management parameters. However, the content is not limited to patent documents. handles information such as text data, image data, audio data, or video data. can be done.
[0007] Content is managed using various meta-information. It is not the content itself, but data that describes the attributes to which the content belongs or related information. For example, the patent number indicates the content of the patent, such as the scope of claims, abstract, The patent number is also associated with meta information (rating information, The patent is managed using meta-information. The numbers are classified according to importance using meta-information. The accuracy and efficiency of the classification are Although it depends on the content of the target document, the difference is likely to be due to the user's experience and skill. There was a need to classify documents, and there was a challenge in improving efficiency.
[0008] Generating a classification model using machine learning requires preparing a large amount of training data. There is a problem that it places an excessive burden on the user. There is a problem that the variation in the number of contents affects the accuracy of the classification model.
[0009] In view of the above problem, one aspect of the present invention is to efficiently generate a classification model and use it to Another object of the present invention is to provide a method for interactively classifying The objective is to provide a graphical user interface for generating classification models. Alternatively, one aspect of the present invention provides a program for classifying information with high probability. One of the challenges is to
[0010] The description of these problems does not preclude the existence of other problems. It is not necessary for one embodiment to solve all of these problems. The subject matter will be self-evident from the description, drawings, claims, etc. It is possible to extract other issues from the drawings, claims, etc. [Means for solving the problem]
[0011] The computer device has a storage device that stores a program. A graphical user interface (hereinafter referred to as G Various information can be displayed via the GUI. Program operation, information assignment, database response, machine learning on computer devices The program can also be used to perform machine learning via a GUI. Calculation results, download educational content or uncategorized content from the database The downloaded content can be displayed on the display device. When we refer to content, we mean educational content, uncategorized content, or categorized content. nothing.
[0012] The proposed content classification system uses machine learning to create a content classification model. A classification model for generated content is used to classify unclassified content. For example, content having multiple pieces of meta information may be used as learning content. The learning label is then added to generate a feature vector from the learning content. When generating a feature vector, meta-information or learning labels are used in the learning context. It can be treated as a feature of the content.
[0013] The training content is treated as training data. The classification model uses the training content The classification model obtained here can be obtained by performing machine learning based on multiple The content with meta information is classified. The classification type is based on the user's purpose. The user can use the classification model to find all It is possible to classify the entire document in less time than it would take to manually or visually judge all documents. It becomes Noh.
[0014] The learning content is downloaded from the learning content stored in the database. or the academic records stored in the storage device of the computer system. Learning content can be used. Learning content includes learning labels. Furthermore, the classification model stored in the database can be downloaded and managed. Alternatively, the classification model may be downloaded from the computer's storage device. A rule may also be used.
[0015] One aspect of the present invention is a method for learning a computer program, comprising: The first feature and learning label are assigned to the content, and the second feature is assigned to the content. A plurality of first classification models are generated by machine learning using a plurality of learning contents. and generating a second classification model using the plurality of first classification models. The second classification model is used to assign judgment information to multiple contents and display them as a graphical user interface. and displaying the content within a user interface.
[0016] One aspect of the present invention is a method for learning a computer program, comprising: The first feature and learning label are assigned to the content, and the second feature is assigned to the content. A plurality of first classification models are generated by machine learning using a plurality of learning contents. a step of generating a plurality of first classification models, a step of calculating an average value from the outputs of the plurality of first classification models, and generating a second classification model using the average value of the number of Adding judgment information to multiple contents and displaying it in a graphical user interface A content classification method including steps.
[0017] One aspect of the present invention is a method for learning a computer program, comprising: The first feature and learning label are assigned to the content, and the second feature is assigned to the content. A plurality of first classification models are generated by machine learning using a plurality of learning contents. and evaluating each of the plurality of first classification models according to a first evaluation criterion. and a step in which a plurality of first classification models are evaluated based on a second evaluation criterion. and a group of evaluation results based on a plurality of first evaluation criteria and a second evaluation criterion. generating a second classification model; and classifying the plurality of contents using the second classification model. and providing the determination information and displaying it within a graphical user interface. It is a method of classifying content.
[0018] In the above configuration, the first evaluation criterion is precision, and the second evaluation criterion is A content classification method based on sensitivity is preferred.
[0019] In each of the above configurations, a step for generating a first classification model using any training content is performed. A content classification method that includes steps is preferred.
[0020] In each of the above configurations, classification information is further provided to the learning content, and the second classification Using the model output, make the same judgment as the classification information from multiple content items that have been given classification labels. A screen for selecting information-bearing content to display within a graphical user interface A content classification method including steps is preferred.
[0021] In each of the above configurations, the feature values given to the learning content or content are management parameters. A content classification method that is meter is preferred.
[0022] In each of the above configurations, the judgment information includes a classification of the content including a classification label or a score. The method is preferred.
[0023] In the above configuration, the graphical user interface is configured to display a specific numerical value of the score. Specify a range and display the corresponding content as a list. A classification method is preferred. [Effects of the Invention]
[0024] One aspect of the present invention can provide a method for accurately classifying information. One aspect of the invention is to provide a user interface that accurately classifies information. Another aspect of the present invention is to provide a program for classifying information with high accuracy. Cut.
[0025] Another aspect of the present invention is an interactive interface for generating a classification model using machine learning. It is possible to provide users with an interface that allows them to prepare training data and evaluate learning results. This reduces the burden on users.
[0026] The effects of one embodiment of the present invention are not limited to the effects listed above. This does not preclude the existence of other effects. Other effects may be affected by this item, as described below. The effects not mentioned in this section are effects that a person skilled in the art would understand without reading the specification or the description. or drawings, etc., and can be extracted appropriately from these descriptions. It should be noted that one aspect of the present invention has at least one of the above-listed effects and / or other effects. Therefore, one aspect of the present invention is, in some cases, There are cases where the effects listed above are not achieved. [Brief explanation of the drawings]
[0027] [Figure 1] FIG. 1 is a flowchart illustrating the classification method. [Figure 2] FIG. 2 is a flowchart illustrating the classification method. [Figure 3] FIG. 3 is a diagram illustrating the connection between classification system 100 and a network. [Figure 4] FIG. 4 is a block diagram illustrating the classification system. [Figure 5] 5A and 5B are diagrams illustrating a graphical user interface. [Figure 6] FIG. 6 is a diagram illustrating a method for generating a classification model. [Figure 7] FIG. 7 is a diagram illustrating a method for generating a classification model. [Figure 8] FIG. 8 is a diagram illustrating a method for generating a classification model. [Figure 9] FIG. 9 is a diagram illustrating a graphical user interface. [Figure 10] FIG. 10 is a diagram illustrating a graphical user interface. DETAILED DESCRIPTION OF THE INVENTION
[0028] In this embodiment, a content classification method will be described with reference to FIGS.
[0029] The content classification method described in this embodiment is based on a program that runs on a computer device. The program is controlled by the memory or the It is stored in the storage or on the network (LAN (Local Area Network) Network), WAN (Wide Area Network), Internet, etc. computers connected via a network, or server computers with databases It is stored in the data.
[0030] The display device of the computer device displays the data and the , and displaying the results of calculations performed on the data by a calculation device included in the computer device. The configuration of the device will be explained in detail with reference to FIG.
[0031] The data displayed on the display device can be displayed in a list format, for example, so that the user can Therefore, the user can easily recognize the operation of the computer through the display device. The interface for easily communicating with the programs on a computer device is called a GUI. and explain.
[0032] The user can use the program's content classification method through the GUI. The user can easily classify the contents using the GUI. In addition, the GUI allows users to visually determine the results of content classification. In addition, users can easily operate the program through the GUI. The content in question is text data, image data, audio data, or video data. It shows information such as:
[0033] Next, we will explain how to classify content using the GUI, following the GUI operation procedure. First, the data processing unit will be described. The data processing unit includes a data collection unit and a data generation unit. For example, the data collection unit can collect data from multiple contents in a database via a GUI. The data generator also receives a file containing the content via the GUI. By assigning learning labels to the Alternatively, learning content with learning labels may be obtained from a database. It should be noted that the plurality of contents are stored in the memory or storage of the computer device. files, or databases, computers, or Or it is data stored on a data server or the like.
[0034] Therefore, the database may contain multiple educational contents or multiple unclassified contents. It is preferable that the contents are listed and saved. The unclassified and unclassified content are given multiple features and learning labels. The label can be modified by the user via the GUI. When the training labels are added to the content, the training content with the training labels is added to the database. can be saved in
[0035] Training content can include validation content that is not labeled for training. The verification content is used to verify the classification model generated using the training content. It can be used for this purpose.
[0036] As an example, let us consider a case where the content is a patent number. As a feature of the issue, multiple pieces of meta information are attached. For example, the meta information may include evaluation information, Number of days passed, number of families, family status, application type, life, pens in family These include the number of patents, number of abandonments within a family, costs, number of inventors, fields, or number of claims. In other words, meta information is a content management parameter. , patent family, or patent family.
[0037] Next, the learning processing unit will be described. The learning processing unit performs classification using learning content. The learning processing unit includes a classification model generating unit or a classification model generating unit. It has a rule evaluation unit.
[0038] The classification model generation unit can generate a classification model. generating a plurality of first classification models by machine learning using the training content; generating a second classification model using the plurality of first classification models. The output values of the first classification model or the second classification model can be displayed in the GUI. The user assigns a label for training the first classification model to the output value (including correction). ) or the user can select new learning content for the output value. can be added.
[0039] The classification model evaluation unit evaluates the classification model generated by the classification model generation unit using the verification content. When the classification model is used to infer the verification content, the classification model is evaluated. The GUI outputs the inference results as judgment information. Judgment information can be added to the item and displayed.
[0040] The user can evaluate the output results of the classification model evaluation unit and, if necessary, change the learning labels. The classification model can be updated by the classification model generation unit. By adding content, the classification model can be updated in the classification model generation unit.
[0041] Next, the judgment processing unit will be explained. The judgment processing unit has a classification inference unit and a list generation unit. For example, the classification inference unit uses the first learning model and the second learning model generated by the classification model generation unit. A learning model is used to infer and classify multiple unclassified pieces of content. The inference results are assigned to each piece of content as judgment information.
[0042] The list generation unit generates a list in a format desired by the user from the content for which the determination information is given. For example, each content can be displayed by country of application. If the classification information is managed by the WHO, the country of application can be used as classification information. In this case, it is preferable to generate a different classification model for each application country. The classification information is not limited to the country of application. For example, one of the meta information of the content may be One of them can be classification information.
[0043] We will explain the case where meta-information is classification information. For example, The patent number may be meta-data such as the patent number of the parent application and the patent number of the divisional application. Information may be provided in the parent application's patent number. A divisional application cannot be filed from the patent number of the parent application, the patent number of the parent application is pending, The patent number is in a state of lapse, a divisional application can be filed from the patent number of the divisional application, A divisional application cannot be filed from the divisional patent number, and the divisional patent number is still in force. Different classification models are generated depending on the status of the patent, the status of the divisional patent number, etc. , can be inferred using the classification model.
[0044] That is, the determination processing unit uses the classification model to infer multiple unclassified contents. The inferred results are added to each content as judgment information and displayed in the GUI. The determination information includes at least a classification label and a score (probability). The GUI also allows you to specify a specific range of scores and view the corresponding content. as a list.
[0045] This example is different from the classification model generation unit described above. The classification model generation unit generates a classification model by using a plurality of learning components. generating a plurality of first classification models by machine learning using the content; calculating an average value from the outputs of a plurality of first classification models; and generating a second classification model using the first classification model or the second classification model. The output values of the classification model 2 can be displayed on the GUI. Alternatively, the user can modify the training labels of the first classification model. The learning content can be added to the output value. It is intended to be calculated using either the arithmetic mean, geometric mean, or harmonic mean. Taste.
[0046] A second classification model is generated using the plurality of average values. Since the output of the first classification model is averaged, noise such as outliers in the training content is eliminated. The influence of noise components can be reduced.
[0047] Next, we will show an example of a classification model generation unit different from the above. generating a plurality of first classification models by machine learning using the training content; evaluating each of the plurality of first classification models according to a first evaluation criterion; evaluating each of the first classification models according to a second evaluation criterion; A second classification model is generated from the evaluation results based on the first evaluation criterion and the evaluation results based on the second evaluation criterion. The output value of the first classification model or the second classification model is The output value can be displayed on the GUI. The user can select the learning value of the first classification model for the output value. The training label can be modified, or the user can assign a training control to the output value. The first evaluation criterion is the accuracy of the confusion matrix, and the second is the accuracy of the The evaluation criterion is the sensitivity of the confusion matrix.
[0048] Evaluating outputs of the plurality of first classification models according to a first evaluation criterion and a second evaluation criterion. The second classification model is generated using the results of the first evaluation criterion. The degree of confusion can be rephrased as the precision of the training labels. The sensitivity of the matrix can be rephrased as the recall for the training labels. The classification model may include the precision and recall of the plurality of first classification models. The second classification model generated using the plurality of first classification models has improved classification accuracy. do.
[0049] In the classification model generation unit described above, for example, m (m represents a natural number) generated The first classification model may be used to generate a second classification model.
[0050] It should be noted that the classification model generation unit has k pieces of learning content (k is a natural number). In this case, the first classification model can generate any number of training contents up to k. In addition, when k pieces of learning content are selected, any learning content can be It can contain two training models, each of which has k different numbers. The contents are sorted by number in groups of q (q is a natural number). A first classification model can be generated by:
[0051] The program can read content from a database and display it in a GUI. The content preferably has meta information in the form of a list. The content will be displayed in accordance with the display format provided by the It is preferable that the meta information be managed in units called records. For example, Each record has an ID (Identification) associated with a number, Content (image data, audio data, or video data) or meta information It is composed.
[0052] In this specification, we focus on meta-information and perform machine learning to generate a classification model. The classification model analyzes meta-information and classifies the feature vectorized content. .
[0053] In addition, in the content classification method described above, learning labels are not used as training data. Classification can be performed using machine learning. For example, classification models include K-means and or DBSCAN (density-based spatial clustering) algorithms such as (g of applications with noise) You can be there.
[0054] The program also generates training content with multiple meta-information and training labels. A classification model can be generated using machine learning. Tree, Naive Bayes, KNN (k Nearest Neighbor), SVM (Su Support Vector Machines), Perceptron, Logistic Regression , or algorithms such as neural networks can be used.
[0055] In addition, the program can switch classification models depending on the number of training contents. For example, when the number of learning contents is small, decision trees, naive Bayes, and Logis If the number of training contents is more than a certain amount, SVM, Random Forest, A neural network may be used. The algorithm uses random forest, which is a decision tree algorithm. a method for selecting data information, a method for selecting training content, or a method for selecting a first classification model; Random sampling or cross-variation can be used to , q items can be selected at a time according to the sorting order of the assigned numbers.
[0056] Next, a content classification method will be explained using the drawings. FIG. 1 shows one embodiment of the present invention. 1 is a flowchart illustrating a content classification method. is controlled by a program running on a computer device. The RAM has a data processing unit, a learning processing unit, or a judgment processing unit to classify content. The program allows users to categorize the content they want through a GUI. In other words, the contents processed by each of the above-mentioned processing units can be Equivalent to Tep.
[0057] In step S11, the user loads a file containing content via the GUI. The file is stored in a database in the data processing unit. The files include educational content, uncategorized content, etc. can be.
[0058] Therefore, the database contains multiple learning contents or multiple unclassified It is preferable that the contents are listed and saved. The user can select the contents displayed in the GUI. You can assign or modify learning labels to learning content that has been registered. The file may contain validation content that has not been labeled for training.
[0059] Step S12 is a learning processing unit that generates a classification model using the loaded file. The generated classification model evaluates the verification content and displays the evaluation results in the GUI. The user can modify the learning labels and the learning content based on the evaluation results. You can instruct additions, etc.
[0060] The user can predict the time-dependent changes in the meta information of the learning content and When a user updates meta information, the classification model can be updated. Therefore, the user can classify the content over time. The classification model can be used to identify content that is expected to grow in value. It is possible to classify groups of content that are expected to decrease in value.
[0061] Step S13 is a determination processing section. The classification model generates inference results for unclassified content. The judgment processing unit can assign judgment information based on the result of the judgment. The content can be displayed in the GUI in the format desired by the user. The GUI also includes at least a classification label and a score. You can specify a range of values and display the corresponding content.
[0062] Next, referring to FIG. 2, the flowchart of FIG. 1 will be explained in more detail. First, step S1 The data processing unit in step S11 performs the data processing in step S21. The data generating unit of step S22 includes a collecting unit.
[0063] The data collection unit in step S21 will be described. The data collection unit in step S21 You can load files from a database. You can also load meta information or content. The meta information can be managed by different databases. This may vary depending on the company, organization, or user. It has the function of collecting meta information about content from different databases. Each database can be located in a different building, in a different region, or even in a different country. Cut.
[0064] Next, the data generation section in step S22 will be described. For example, each record can be managed by The code contains the ID associated with the number, the content (image data, audio data, or video data), The user can select the commands displayed in the GUI. By assigning learning labels to content, learning content can be generated. Cut.
[0065] Next, the details of step S12 will be described. A classification model generation unit in step S23, a classification model evaluation unit in step S24, and The process includes an output result determination process in S25.
[0066] The classification model generation unit in step S23 will be described. The classification model generation unit generates a classification model for multiple learning contents. A plurality of first classification models can be generated by machine learning using the above. The first classification model can be used to generate a second classification model. The output values of the classification model or the second classification model can be displayed.
[0067] After the output result determination process in step S25 described later, the user can Labels for training the first classification model can be added (including corrections). The user can add new learning content for that output value. It is possible to predict changes over time in the meta information of the content and update the meta information. For the effect obtained by updating the meta information by the user, see the explanation of step S12. You can pour drinks.
[0068] Next, the classification model evaluation unit in step S24 will be described. By using the content for use, the classification model generated by the classification model generation unit is evaluated. The classification model outputs the results of inferring the verification content as judgment information. The GUI can display each evaluation content with its evaluation information attached. .
[0069] Next, the output result determination process of step S25 will be described. The output result of the classification model evaluation part of step S24 is judged, and the classification model of the content is sufficiently trained. The user can determine that the classification model generation is complete (OK) through the GUI. For example, if the user determines that the content classification model is not sufficiently trained (NG), The user can return to step S23 to change the learning label and the learning content. The classification model is updated by adding new content or updating meta information.
[0070] Next, the details of step S13 will be described. The system has a classification inference unit in step S26 and a list creation unit in step S27.
[0071] The classification inference unit in step S26 will be described. Inferring multiple unclassified contents using the first and second learning models The classification inference unit receives the unclassified data generated by the data generation unit in step S22. Content is given, and the classification model determines the inference result for each content. It is assigned as fixed information.
[0072] The list creation unit in step S27 will be described. The obtained content can be listed in the format desired by the user and displayed in a GUI. Each content may be given classification information different from the meta information. For example, if classification information is given in the training content, the generated classification model can be A different classification model can be generated for each type of information. One of the meta information can be classification information.
[0073] The judgment information includes at least a classification label and a score. I can specify a specific range of scores and display the corresponding content in a GUI. can be displayed in.
[0074] FIG. 3 shows a classification system 100 having the above-described content classification method and a network FIG. 1 is a diagram illustrating a connection with a network.
[0075] The classification system 100 is connected to a communication network LAN1. database DB1, or client computers CL1 to CLn (n is a natural number) The communication network LAN1 is connected to the communication network LAN2 via a network. The network can be the Internet, a communication network WAN, or The communication network LAN2 can use satellite communication. Client computers CL11 to CL1n are connected.
[0076] The classification system 100 includes a database DB1, a database DB2, a client computer The data is recorded on the client computers CL1 to CLn or the client computers CL11 to CL1n. Content generation and content classification are performed using files containing stored content. ,model generation, and unclassified content can be classified.
[0077] In addition, the user can input instructions to the GUI from a program running on the classification system 100. For example, a user can send a message to a server located in a different country via the Internet. The information in the database is used to generate the classification model described above, and to classify the unclassified content. This means that content or meta-information can be shared across different databases or clusters. The information may be stored in the client computer.
[0078] The GUI is connected to the database DB1, the database DB2, and the client computer. the client computers CL1 to CLn, or the client computers CL11 to CL1n. The classification results stored in the storage device of the computer device are displayed. It is possible.
[0079] FIG. 4 is a block diagram illustrating the classification system 100 described in FIG. The system 100 includes a GUI (Graphical User Interface) 110, The GUI 110 includes an input unit 111 and an output unit 120. The input unit 111 has a function of selecting the content load source and a function of learning. The output unit 112 has a function to input practice labels. The function to display the content list generated by the classification model and the function to display the judgment information output by the classification model are also available. The meta information included in the displayed content can be modified by the user via the GUI. It can be corrected.
[0080] The calculation unit 120 includes a data processing unit 121, a learning processing unit 122, and a determination processing unit 123. The data processing unit 121 includes a data collection unit and a data generation unit. The learning processing unit 122 includes a classification model generation unit that creates a classification model, and a classification unit that classifies the classification model. The classification model evaluation unit has a classification model evaluation section that evaluates the classification model. The judgment processing unit 123 has a function of performing an evaluation result judgment process. and an output list creation unit that lists the results of classification by the classification inference unit. The unit 120 is a unit in which a program stored in a storage unit of a computer device is executed by a microprocessor. However, the program is processed using a DSP (Digital Signal Processor). l Processor), or GPU (Graphics Processing Unit) Calculations can be performed using the .
[0081] The storage unit 130 stores content and meta information that have been generated by loading from a database or the like. The information is listed and temporarily stored.
[0082] The storage unit 130 includes, for example, 1T (transistor) 1C (capacitor) type memory cells. DRAM (Dynamic Random Access Memory) can be used. The transistors used in the memory cells of M may be OS transistors. A transistor is a transistor that has a metal oxide in the semiconductor layer. A memory device that uses transistors is called an "OS memory." For example, RAM with 1T1C type memory cells is called "DOSRAM (Dyna The new RAM is called "micronic oxide semiconductor RAM."
[0083] The off-state current of the OS transistor is very small. Therefore, the DOSRAM is This reduces the frequency of refresh operations, thereby reducing the power required for refresh operations. The off-state current is the current that flows between the source and drain when the transistor is in the off state. When the transistor is an n-channel type, for example, the threshold voltage is about 0 V to 2 V. If so, the current flow between the source and drain when the voltage between the gate and source is negative The current that flows can be called the off-state current.
[0084] 5A is a diagram illustrating the configuration of the GUI 30. As an example, the GUI 30 has p This shows the management screen that displays a list of learning content. Learning content is organized by record. The record is managed as follows: number (No) 31, content (ID) 32, feature quantity Meta information (Feature) 33 (meta information (F1) 33a to meta information (Fm) 33m), classification information (Case) 34 (classification information (C1) 34a to classification information (Cq) 3 4q), and training labels (J-Label) 35. For example, in Figure 5A, The learning label 35 gives either of the two values "Yes" or "No." The label 35 is not limited to being binary, but may be ternary or more.
[0085] FIG. 5B is a diagram illustrating the configuration of the GUI 30A. The GUI 30A is A management screen is created that infers n unclassified contents and displays the inferred judgment information in a list. The unclassified content is number 31, content 32, as well as the learning content. It has meta information 33 and classification information 34. Furthermore, each record has a classification label ( A-Label 36 and score 37 are assigned as judgment information.
[0086] The GUI 30 and GUI 30A can be managed on the same display screen. In Figure 9 or Figure 10, the learning content and the assessment information are displayed on the same management screen. An example of the GUI display that can do this is shown below.
[0087] Figure 6 shows the complex data associated with the above-mentioned learning content Sample through machine learning. FIG. 1 is a diagram illustrating a method for generating a classification model using numerical features. Each of the metadata is used to manage the content. In this embodiment, the calculation unit F, the calculation unit S, and the calculation unit V , the first classification model, and the second classification model are used to explain how to generate a classification model. .
[0088] The study content Sample(1) to the study content Sample(k) are Each of them is annotated with j features and a label for learning. For example, the calculation unit F1 calculates the learning content Sample(1) by The feature vector Vlabel1(1) can be generated in the format Fk is a feature of the learning content Sample(k) in a format that can be processed by a computer. The vector Vlabel1(k) can be generated. l1(1) is generated by the calculation unit F1 by giving different weighting coefficients to each feature. In addition, the feature vector Vlabel1(1) is randomly selected. It can be generated using up to j features.
[0089] Next, a plurality of first classification models are generated. As an example, the calculation unit S1 calculates the feature vector Vlabel 1(1) to generate a first classification model using the feature vector Vlabel1(k). The number of feature vectors Vlabel1 given to the calculation unit S1 is k or less. As a different example, the calculation unit Sm may be configured to calculate different feature vectors Vlabel1(1) to Vlabel2(2). The feature vector Vlabel1(k) can be used to generate a first classification model. Therefore, two different first classification models are each based on k or less feature vectors Vlabe l1 and any one of the feature vectors Vlabel1 contains the same feature vector It is possible.
[0090] The k training content samples selected to generate the first classification model are The learning content may be selected randomly or sorted by the number assigned to the learning content. If the training content is selected randomly, the first classification model The learning content can include variations in the learning content. If the selection is made in the order sorted by the number, either the chronological order or the meta information It may include a trend along with a number assigned based on the symptom.
[0091] Therefore, the first classification model is The feature vector Vlabel1 generated from the content Sample(k) is used to generate the feature vector. vector Vlabel2 can be generated.
[0092] The second classification model is generated by the calculation unit V1. For example, the calculation unit V1 generates m generating a second classification model using the feature vectors Vlabel2; The second classification model uses the feature vector Vlabel2(1) to the feature vector Vl Using abel2(m), we can generate classification models with different features. .
[0093] Therefore, the second classification model is The output value is calculated using the feature vector Vlabel1 generated from the content Sample(k). The GUI can display the output value POUT. The output value POUT includes a classification label, which is judgment information, and a score. Therefore, the second classification model can classify the content. The rule can assign judgment information to each content.
[0094] To make inference using the classification model described in Figure 6, the learning context for the classification model is The judgment result is obtained by giving unclassified content to the Content Sample. Unlike the learning content, the content is not labeled for learning purposes.
[0095] FIG. 7 is a diagram illustrating a method for generating a classification model different from that shown in FIG. 6. In FIG. 7, and explain the differences between the invention and the embodiment. In this case, the same reference numerals are used in common between different drawings for parts having similar functions, and the repeated reference numerals are used in different drawings. The explanation will be omitted.
[0096] In Figure 7, the average value Av of m feature vectors Vlabel2 is calculated, and the feature vector V The second classification model generates p feature vectors Vlabel_a The second classification model can be generated using m feature vectors Vlabel By calculating the average value Av of 2, it is possible to generate a classification model with different features. The generated classification model can accurately classify the content.
[0097] FIG. 8 is a diagram illustrating a method for generating a classification model different from that shown in FIG. 7. 7, and explain the differences between the invention and the embodiment. The same reference numerals are used in different drawings for parts having similar functions, and their repetition The explanation will be omitted.
[0098] In Fig. 8, the evaluation criteria for evaluating m feature vectors Vlabel2 are For example, the evaluation and judgment unit JG1 is provided with precision (Pre cision) is given and each feature vector Vlabel2(1) is evaluated. Next, the evaluation and judgment unit JG1 determines the sensitivity as the second evaluation criterion. ty) and evaluate each feature vector Vlabel2(1) The evaluation unit JG1 outputs the evaluation result Vlabel_b(1).
[0099] The second classification model is the evaluation result Vlabel_b(1) to the evaluation result Vlabel_b For example, multiple feature vectors Vlabel2 are generated using different first may be evaluated by one criterion and a second criterion, or may be evaluated by the same criterion. Although not shown in FIG. 8, the first evaluation criterion and the second evaluation criterion may be used in the same manner as in FIG. The average value of the evaluation results Vlabel_b based on the second evaluation criterion can be calculated.
[0100] The second classification model uses the evaluation results of m feature vectors Vlabel2, It is possible to generate classification models with different features. This allows for accurate classification of content.
[0101] FIG. 9 is a diagram illustrating the GUI 50. The GUI 50 displays content (learning content) Display areas for classified content, unclassified content, and classified content, and files containing the content Icon 58a to select the download source, the address where the selected file will be saved A text box 58b for displaying information, and an icon for performing machine learning (Learning g Start)59.
[0102] The display area shows an example where eight records are loaded. Each record has a number (No) 51, ID (Index) 52, and feature ) 53, classification information (Case) 54, learning label (JL) 55, classification label (AL) 5 6 and score (Prob) 57. The feature 53 is used as detailed information. The feature quantities F(1) 53a to F(j) 53j can be displayed. The classification information 54 is a natural number. Classification information C(4) 54d can be displayed. Note that classification information can be expressed as natural numbers. It can have a variety of types.
[0103] In addition, Figure 9 shows how the training content and unclassified content are classified by the classification model. 1 shows an example of displaying the results on GUI 50.
[0104] As an example, record numbers No. 1 to No. 3 correspond to study content. The contents for learning are given learning labels, and the record numbers No. 1 to No. 3 are , classification information is given.
[0105] Record numbers No. 4 to No. 8 correspond to classified content. Each content is given a classification label 56 and a score 57. The classification model obtained by learning record number 1 and record number 3 is used to classify records. The classification results for record numbers No. 4 to No. 7 are shown. Using the classification model obtained by learning code number No. 2, record number No. 8 The classification results are displayed. Due to space limitations, only up to 8 records are displayed in Figure 9. However, the number of records can be of multiple types.
[0106] However, when dealing with a large number of records, there are display issues. Therefore, classification label 5 6 or score 57 preferably has a sorting function. As an example of a sorting condition, G The UI can select and display the determination result for which the classification label 56 is "Yes." The GUI can be used to specify a numerical range for the score 57. When sorting conditions are given, the GUI uses the same characteristics as the training content given the training data. Content having the signature can be classified and displayed.
[0107] For example, let us consider the case where the content is a patent number. The patent number contains multiple meta-information. For patent numbers where the patent rights are maintained, the learning label will include For patent numbers that are abandoned, the learning label should read: Give it "No." Then run machine learning to generate a classification model.
[0108] The classification model described above can assign judgment information to unclassified content. The judgment information includes a classification label 56 and a score 57. For example, the user can Using the sort function, the classification label 56 is assigned a "No." Furthermore, the score 57 is assigned a "0 By providing the above sorting conditions to the GUI, Select records with similar characteristics to the abandoned patent number educational content. It can be displayed as follows.
[0109] FIG. 10 is a diagram illustrating a GUI 50A different from that shown in FIG. 9. This is an example of an efficient GUI display when handling large amounts of data. Note that Figure 10 differs from Figure 9 in the following ways: The same parts or similar functions are described in the structure of the invention (or the structure of the embodiment). The same reference numerals are used in common among different drawings for parts having the same meaning, and repeated explanations thereof will be omitted. do.
[0110] In Figure 10, records related to any selected classification information can be categorized and displayed. 9. In FIG. 10, the types of classification information C(1) to C(4) are The display can be switched depending on the type.
[0111] The user can view multiple features53 assigned to a record and the classification model's judgment information. If the classification accuracy is sufficient, the classification model update is completed. If the classification accuracy is not sufficient, the user-specified label is used. To assign a learning label to a record that does not have a label assigned, click icon 59. The classification model can be updated by using the following. The user may predict changes over time and update the feature quantity 53. ,The classification model can include changes in the classification model over time.,Therefore, the user The classification model can be used to obtain the change in classification of content over time. A group of content that is expected to grow or decrease in value You will be able to classify them.
[0112] Although not shown, the feature amount 53, classification information 54a to 54d, and learning labels 5 5. The numerical values and label information included in the classification label 56 or score 57 are displayed in the order You can change the order or filter the selected values and label information using the filter function. This allows users to sort and display the data in the order they need. The judgment results can be efficiently evaluated.
[0113] The content classification method described with reference to FIGS. 1 to 10 classifies information with high probability. For example, the GUI is suitable for classifying high probability information. The program is based on the fact that new training data (training labels) are given to the classification model. The classification model can be updated by It is possible to classify information with a high degree of accuracy.
[0114] Furthermore, the generated classification model can be stored in the device itself or in external memory. , can be called and used when classifying new files. The classification model can be updated according to the method described above while adding
[0115] As described above, the structures and methods described in this embodiment mode may be appropriately combined with the structures and methods described in other embodiments. They can be used in combination. [Explanation of symbols]
[0116] CL1: Client computer, CL1n: Client computer, CL11: Client computer, CLn: Client computer, DB1: Database ,DB2: Database, LAN1: Communication network, LAN2: Communication network, Vlabel1: Features Vector, Vlabel2: feature vector, 31: number, 32: content, 33: meta Information, 34: Classification information, 35: Learning labels, 50: GUI, 50A: GUI, 51: Number No., 53: Feature, 54: Classification information, 56: Classification label, 57: Score, 58a: Icon ,58b:Text box, 59:Icon, 100:Classification system, 110:GU I, 111: input unit, 112: output unit, 120: calculation unit, 121: data processing unit, 122 : learning processing unit, 123: judgment processing unit, 130: storage unit
Claims
[Claim 1] A learning content and a content, a first feature amount and a learning label are assigned to the learning content; a second feature is assigned to the content; generating a plurality of first classification models by machine learning using a plurality of the training contents; generating a second classification model using the plurality of first classification models; assigning determination information to the plurality of pieces of content using the second classification model and displaying the determination information in a graphical user interface; How to categorize content, including:
Citation Information
Patent Citations
Machine learning approach to determining document relevance for searching over large electronic collections of documents
JP2009104630A