Information processing apparatus

By extracting keywords from whole documents, text blocks, and images, identifying non-redundant keywords, and excluding related classification labels, the error detection problem of LLM models in document classification is solved, thus improving the accuracy of classification.

CN122286402APending Publication Date: 2026-06-26TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2025-12-22
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, when using models trained with Large Language Models (LLM) to classify document data, it is difficult to automatically detect misclassifications, resulting in a high risk of classification errors.

Method used

By acquiring document data, keywords are extracted from the whole document, text blocks, and images. Non-redundant keywords are identified, and based on the difference set determination, classification labels related to non-redundant keywords are excluded to classify the document data.

Benefits of technology

It effectively suppresses misclassification by judging the importance of keywords from multiple perspectives, eliminating inappropriate category labels, and improving the accuracy of document data classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122286402A_ABST
    Figure CN122286402A_ABST
Patent Text Reader

Abstract

This disclosure relates to an information processing apparatus. The information processing apparatus includes an acquisition unit for acquiring document data, a first extraction unit for extracting a first keyword from the entire document data, a second extraction unit for extracting a second keyword from text boxes included in the document data, a determination unit for finding the difference between the first keyword and the second keyword and determining non-redundant keywords based on the difference, and a classification unit for classifying the document data by assigning classification labels to the document data. The classification unit excludes classification labels related to the non-redundant keywords from candidates for classification labels to be assigned to the document data, and then assigns the classification labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the technical field of information processing devices. Background Technology

[0002] As an example of this type of system, a system has been proposed in which query data is generated based on documents using a large language model (LLM), and a retrieval model for a chatbot is trained using document and query data pairs (see Japanese Unexamined Patent Application Publication No. 2023-076413 (JP 2023-076413 A)). Summary of the Invention

[0003] When classifying document data using models trained through machine learning methods such as LLM, there is a risk of performing incorrect classifications. However, it is difficult to automatically detect misclassifications that have already been performed.

[0004] This disclosure has been made with the above-mentioned problems in mind, and its purpose is to provide an information processing apparatus that can appropriately classify document data.

[0005] The information processing apparatus according to the present disclosure includes: an acquisition unit for acquiring document data; a first extraction unit for extracting a first keyword from the entire document data; a second extraction unit for extracting a second keyword from text boxes included in the document data; a determination unit for finding the difference between the first keyword and the second keyword and determining non-redundant keywords based on the difference; and a classification unit for classifying the document data by assigning classification labels to the document data, wherein the classification unit excludes classification labels related to the non-redundant keywords from candidates of classification labels to be assigned to the document data, and then assigns the classification labels. Attached Figure Description

[0006] The features, advantages, and technical and industrial significance of exemplary embodiments of the present invention will now be described with reference to the accompanying drawings, wherein like symbols denote like elements, and wherein: Figure 1 This is a block diagram illustrating the hardware configuration of an information processing apparatus according to an embodiment; Figure 2 This is a block diagram illustrating the functional configuration of an information processing apparatus according to an embodiment; Figure 3 This is a flowchart illustrating the operation of an information processing apparatus according to an embodiment; and Figure 4 This is a schematic diagram illustrating an example of an operation used to extract keywords. Detailed Implementation

[0007] Embodiments of the information processing apparatus will now be described with reference to the accompanying drawings.

[0008] Hardware configuration First, refer to Figure 1 The hardware configuration of the information processing apparatus according to an embodiment is described. Figure 1 This is a block diagram illustrating the hardware configuration of an information processing apparatus according to an embodiment.

[0009] exist Figure 1 In this embodiment, the information processing apparatus 10 comprises a computing device 110, a storage device 120, a communication device 130, an input device 140, and an output device 150. The computing device 110, storage device 120, communication device 130, input device 140, and output device 150 are interconnected via a data bus.

[0010] The computing device 110 is configured to perform various types of computational processing within the information processing device 10. The computing device 110 may include a processor. The computing device 110 may have a single processor or multiple processors. That is, the computing device 110 may have more than one processor. Note that the processor may be a multi-core processor. When the computing device 110 has a single processor that functions as a multi-core processor, the computing device 110 can be considered to logically have multiple processors.

[0011] The processor of the computing device 110 may be at least one of, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), and a tensor processing unit (TPU).

[0012] Storage device 120 can be at least one of, for example, random access memory (RAM), read-only memory (ROM), hard disk drive, magneto-optical disk drive, solid-state drive (SSD), and optical disk array. That is, storage device 120 can be implemented using a single device or using multiple devices.

[0013] Storage device 120 is capable of storing desired data. Storage device 120 can store computer program CP executed by computing device 110. While computing device 110 is executing computer program CP, storage device 120 can temporarily store data temporarily used by computing device 110.

[0014] Note that the computer program CP can be recorded on a computer-readable and non-transitory recording medium. In this case, the computer program CP can be stored in the storage device 120 by reading the recording medium using a recording medium reading device (omitted from the figure) included in the information processing apparatus 10. Note that at least one of optical discs, magnetic media, magneto-optical discs, semiconductor memory, and any other medium capable of storing programs can be used as the aforementioned recording medium. The computer program CP can be obtained from a device outside the information processing apparatus 10 (omitted from the figure) via a communication device 130. In other words, the computer program CP can be downloaded from an external device to the storage device 120 of the information processing apparatus 10.

[0015] The computing device 110 (e.g., a processor) together with the storage device 120 storing the computer program CP (i.e., together with the storage device 120 and the computer program CP stored in the storage device 120) can perform the processing to be performed by the information processing device 10. For example, logical function blocks for performing the processing to be performed by the information processing device 10 can be implemented by the computing device 110 executing the computer program CP within the computing device 110 (e.g., within the processor).

[0016] The communication device 130 is configured to communicate with a device external to the information processing device 10. Note that the communication device 130 can perform wired communication or wireless communication.

[0017] Input device 140 is a device capable of receiving information input to information processing device 10 from an external source. Input device 140 may include operating devices (e.g., keyboard, mouse, touch panel, etc.) operable by a user of information processing device 10. Input device 140 may include, for example, a recording medium reading device capable of reading information recorded on a recording medium such as a Universal Serial Bus (USB) memory, which can be attached to and detached from information processing device 10. Note that when information is input to information processing device 10 via communication device 130 (i.e., when information processing device 10 obtains information via communication device 130), communication device 130 may be used as an input device.

[0018] Output device 150 is a device capable of outputting information from information processing device 10 to an external device. Output device 150 may have a display device capable of outputting visual information (such as text, images, etc.) as such information. Output device 150 may have a speaker capable of outputting auditory information (such as sound) as such information. Output device 150 may be configured to output such information (e.g., control information for other devices, etc.) to other devices. Output device 150 may be capable of outputting information to a recording medium (such as, for example, a USB memory) that can be attached to and detached from information processing device 10. Note that when information processing device 10 outputs information via communication device 130, communication device 130 may be used as an output device.

[0019] Functional Configuration Next, we will refer to Figure 2 The functional configuration of the information processing apparatus 10 according to an embodiment is described. Figure 2 This is a block diagram illustrating the functional configuration of an information processing apparatus according to an embodiment.

[0020] exist Figure 2 In this embodiment, the information processing device 10 is configured to classify document data and store the document data in the document data database 300. The information processing device 10 includes a document data acquisition unit 210, a first extraction unit 220, a second extraction unit 230, a third extraction unit 240, a determination unit 250, an alarm output unit 260, and a classification unit 270 as components for implementing its functions. Note that each of the document data acquisition unit 210, the first extraction unit 220, the second extraction unit 230, the third extraction unit 240, the determination unit 250, the alarm output unit 260, and the classification unit 270 can be a processing block implemented by the aforementioned computing device 110.

[0021] The document data acquisition unit 210 is configured to acquire document data. Document data includes text, and examples include technical documents, specifications, etc. Document data can also include images in addition to text. For example, the document data acquisition unit 210 can acquire document data input by a user. Alternatively, the document data acquisition unit 210 can acquire document data pre-stored in a storage device or the like.

[0022] The first extraction unit 220 is configured to extract a first keyword from the overall document data. The first keyword can be a keyword that is very important when viewing the overall document data. The first keyword can be, for example, a frequently occurring keyword throughout the entire document data. The first extraction unit 220 can extract multiple first keywords from a single document.

[0023] The second extraction unit 230 is configured to extract second keywords from text boxes included in document data. The second keyword can be a keyword that is very important within the text block. For example, the second keyword can be a keyword used in the title of the text block. It should be noted that the term "text block" here refers to a set of sentences (a collection of sentences (lump)) in the document data, and is not limited to sentences divided into blocks. For example, a text block can be a paragraph. The second extraction unit 230 can extract multiple second keywords from a single text block.

[0024] The third extraction unit 240 is configured to extract third keywords from images included in document data. The third keywords can be extracted from text contained within the image. Alternatively, the third keywords can be extracted from a summary of the image. Alternatively, the third keywords can be extracted by identifying objects in the image. The third extraction unit 240 can extract multiple third keywords from a single image.

[0025] The determination unit 250 determines the non-redundant keywords among the first keyword extracted by the first extraction unit 220, the second keyword extracted by the second extraction unit 230, and the third keyword extracted by the third extraction unit 240. More specifically, the determination unit 250 finds the difference set relative to the first keyword, the second keyword, and the third keyword, and determines the non-redundant keywords based on the difference set.

[0026] Alarm output unit 260 calculates the proportion (hereinafter appropriately referred to as "non-redundancy ratio") of the non-repeating keywords determined by judgment unit 250 among all keywords extracted by first extraction unit 220, second extraction unit 230, and third extraction unit 240. Alarm output unit 260 is configured to output alarm information when the non-redundancy ratio exceeds a predetermined threshold. The alarm information may be information output to a user. For example, the alarm information may be displayed as text or an image on a monitor, etc. Alternatively, the alarm information may be sound output from a speaker, etc. The alarm information may be a notification that a classification error may exist in the document data being processed. The alarm information may be, for example, a prompt to perform a manual classification check.

[0027] Classification unit 270 is configured to classify document data by assigning classification labels to the document data. For example, classification unit 270 selects and assigns suitable classification labels to the document data from a plurality of pre-prepared classification labels. Note that classification unit 270 can perform classification using methods other than assigning classification labels. Classification unit 270 can classify the document data using a model trained via machine learning. This model can be a model that takes document data as input and outputs classification labels to be assigned to the document data. Classification unit 270 can be configured to store the classified document data in a document data database.

[0028] Specifically, classification unit 270 excludes classification labels related to non-redundant keywords determined by decision unit 250 from the candidate classification labels to be assigned to document data, and then assigns classification labels. Therefore, classification labels related to non-redundant keywords are not assigned to document data.

[0029] Document data database 300 is configured to store document data categorized by classification unit 270. The document data stored in document data database 300 can be appropriately readable. For example, the document data stored in document data database 300 can be configured to be searchable using category tags as criteria.

[0030] Note that, although Figure 2 The figure illustrates an example of providing a document data database 300 from outside the information processing device 10, but the information processing device 10 can be configured to include the document data database 300.

[0031] Operation process Next, we will refer to Figure 3 and Figure 4 The operation flow of the information processing apparatus 10 according to an embodiment is described. Figure 3 This is a flowchart illustrating the operation of an information processing apparatus according to an embodiment. Figure 4 This is a schematic diagram illustrating an example of an operation used to extract keywords.

[0032] like Figure 3 As shown, when the operation of the information processing apparatus 10 according to the embodiment begins, the document data acquisition unit 210 first acquires document data (step S101). Then, the first extraction unit 220 extracts a first keyword (step S102). The second extraction unit extracts a second keyword (step S103). The third extraction unit extracts a third keyword (step S104).

[0033] like Figure 4As shown, the first extraction unit 220 extracts the first keyword from the entire document data. The second extraction unit 230 extracts the second keyword from the text blocks included in the document data. The third extraction unit 240 extracts the third keyword from the images included in the document data. Note that there is no particular restriction on the order in which the first, second, and third keywords are extracted. That is, the processes in steps S102, S103, and S104 can be executed consecutively or simultaneously in parallel.

[0034] return Figure 3 Once the keywords have been extracted, the decision unit 250 finds the difference between the first keyword, the second keyword, and the third keyword, and determines whether any non-redundant keywords exist (step S105). Note that if there are no non-redundant keywords ("No" in step S105), the document data is classified as is (step S109), and the series of operations ends.

[0035] On the other hand, when non-redundant keywords exist ("Yes" in step S105), the classification unit 270 excludes classification labels related to non-redundant keywords from the candidates of classification labels to be assigned to document data (step S106).

[0036] The alarm output unit 260 further determines whether the proportion of non-redundant keywords exceeds a predetermined threshold (step S107). If the proportion of non-redundant keywords exceeds the predetermined threshold ("Yes" in step S107), the alarm output unit 260 outputs alarm information (step S108). On the other hand, if the proportion of non-redundant keywords does not exceed the predetermined threshold ("No" in step S107), step S108 is skipped. That is, the alarm output unit 260 does not output alarm information.

[0037] Subsequently, classification unit 270 classifies the document data (step S109). Here, in step S106, classification labels related to non-redundant keywords are excluded from the candidates for classification labels to be assigned to the document data. Therefore, classification unit 270 selects classification labels to be assigned to the document data from the unexcluded classification labels.

[0038] Technical effect Next, the technical effects obtained by the information processing apparatus 10 according to the embodiment will be described.

[0039] For reference Figures 1 to 4According to the embodiment, the information processing apparatus 10 determines non-redundant keywords from keywords extracted from each of the document data as a whole, text blocks, and images. Then, classification labels related to non-redundant keywords are excluded from the candidates, and the document data is classified. Therefore, the assignment of classification labels related to keywords that are not closely related to the document data can be suppressed. That is, incorrect classification can be suppressed. In particular, according to this embodiment, keywords extracted from each of the document data as a whole, text blocks, and images are used, so the importance of keywords can be determined from various perspectives, and inappropriate classification labels can be excluded in advance.

[0040] The present invention, derived from the above embodiments, will now be described.

[0041] The information processing apparatus according to the present disclosure includes: an acquisition unit for acquiring document data; a first extraction unit for extracting a first keyword from the entire document data; a second extraction unit for extracting a second keyword from text boxes included in the document data; a determination unit for finding the difference between the first keyword and the second keyword and determining a non-redundant keyword based on the difference; and a classification unit for classifying the document data by assigning classification labels to the document data, wherein the classification unit excludes classification labels related to the non-redundant keyword from candidates of classification labels to be assigned to the document data, and then assigns the classification labels. In the above embodiments, "document data acquisition unit 210" corresponds to an example of "acquisition unit", "first extraction unit 220" corresponds to an example of "first extraction unit", "second extraction unit 230" corresponds to an example of "second extraction unit", "determination unit 250" corresponds to an example of "determination unit", and "classification unit 270" corresponds to an example of "classification unit".

[0042] The information processing apparatus according to the above scheme may further include a third extraction unit, which extracts a third keyword from an image included in the document data, wherein the determination unit finds the difference set between the first keyword, the second keyword, and the third keyword, and determines the non-redundant keyword based on the difference set. Therefore, non-redundant keywords can be determined while considering the image included in the document data. In the above embodiments, "third extraction unit 240" corresponds to an example of "third extraction unit".

[0043] The information processing device according to the above scheme may further include an alarm unit, which outputs an alarm message when the ratio of the non-redundant keywords to all extracted keywords exceeds a predetermined threshold. Therefore, the user can be notified that the ratio of non-redundant keywords is high. In this case, for example, a manual check can be performed. In the above embodiment, "alarm output unit 260" corresponds to an example of "alarm unit".

[0044] This disclosure is not limited to the above embodiments, and various modifications may be made appropriately without departing from the spirit or essence of the invention as can be discerned from the claims and description as a whole, and information processing apparatuses incorporating such modifications also fall within the technical scope of this invention.

Claims

1. An information processing apparatus, comprising: The acquisition unit retrieves document data; The first extraction unit extracts the first keyword from the entire document data; The second extraction unit extracts the second keyword from the text box included in the document data; The determination unit finds the difference between the first keyword and the second keyword, and determines the non-redundant keyword based on the difference. as well as A classification unit that classifies the document data by assigning classification labels to the document data, wherein... The classification unit excludes classification tags associated with the non-redundant keywords from the candidates of classification tags to be assigned to the document data, and then assigns the classification tags.

2. The information processing apparatus according to claim 1 further includes a third extraction unit, the third extraction unit extracting a third keyword from an image included in the document data, wherein the determination unit finds the difference set between the first keyword, the second keyword and the third keyword, and determines the non-redundant keyword based on the difference set.

3. The information processing apparatus according to claim 1 or 2 further includes an alarm unit, which outputs an alarm message when the ratio of the non-redundant keywords to all extracted keywords exceeds a predetermined threshold.

Citation Information

Patent Citations

  • Method, computer device, and computer program for providing dialogue dedicated to domain by using language model

    JP2023076413A