Information Processing Apparatus, Information Processing Method, and Information Processing Program

The information processing apparatus addresses the challenge of inaccurate third-party annotation in machine learning models by determining the attribute and provisional labels of conversation data and requesting annotation from relevant participants, ensuring accurate and reliable annotation.

JP7694495B2Active Publication Date: 2025-06-18TOYOTA JIDOSHA KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022117576
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-06-18
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

When annotating a machine learning model for estimating the state of a conversation participant from conversation utterance data, data for a predetermined annotation candidate may not gather, and annotations by a third party who has not participated in the actual conversation may be inaccurate, leading to data shortages.

Method used

An information processing apparatus and method that determine the attribute of utterance data, compare it with pre-generated insufficient data, and use a model for pre-label determination to identify provisional labels. The apparatus then requests annotation from the appropriate conversation participant based on participant information, ensuring accurate annotation.

Benefits of technology

This approach allows for accurate annotation of conversation data compared to third-party annotation, reducing data shortages and improving the reliability of machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694495000001
    Figure 0007694495000001
  • Figure 0007694495000002
    Figure 0007694495000002
  • Figure 0007694495000003
    Figure 0007694495000003
Patent Text Reader

Abstract

To provide an information processing device and information processing program, which enable more accurate annotation of conversation data compared with a case where a third party annotates the conversation data.SOLUTION: An information processing device disclosed herein is configured to: identify speakers of the utterance data (100) to determine whether the utterance data has an insufficiency attribute or not (104); identify a label of the utterance data (114) to determine whether the utterance data has an insufficiency label or not (116); and determine a delivery destination to generate a message according to the delivery destination, and deliver the message to the delivery destination determined to be a request destination for annotation of the utterance data with the insufficiency attribute and the insufficiency label (126-136).SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus , Information processing method, and an information processing program.

Background Art

[0002] Patent Document 1 proposes an information processing apparatus including storage means for storing, for each of a plurality of learning data, a plurality of pieces of label information associated with the learning data and attribute information corresponding to the label, and label determination means for determining a label for the target learning data among the plurality of learning data based on the plurality of labels corresponding to the target learning data and the attribute information corresponding to the label.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When annotating a machine learning model for estimating the state of a conversation participant from conversation utterance data, it is conceivable that data for a predetermined annotation candidate does not gather.

[0005] In addition, since the label is determined by the subjectivity of the individual user, annotation by a third party who has not participated in the actual conversation cannot accurately perform annotation, which may promote data shortage.

[0006] The present invention has been made in consideration of the above facts, and an object of the present invention is to provide an information processing apparatus and an information processing program capable of accurately performing annotation as compared with the case where a third party annotates conversation data. , Information processing method, and an information processing program.

Means for Solving the Problem

[0007] When performing annotation of a machine learning model that estimates the state of a party to the conversation from the utterance data of the conversation, the information processing apparatus according to the first aspect A data attribute determination unit that determines the attribute of the utterance data of the conversation from the utterance data, and collates the determined attribute with pre-generated insufficient data to determine whether the utterance data corresponds to the insufficient data; using a model for pre-label determination to determine a provisional label for the utterance data corresponding to the insufficient data, and comparing the determined provisional label with pre-determined insufficient label information to determines whether the utterance data is insufficient ru a a determination unit that determines whether the utterance data is annotation candidate utterance data, and when it is determined by the determination unit that the utterance data is the annotation candidate utterance data that is insufficient, Based on the participant information of the conversation, determine the destination of the annotation request, at least one of the speaker and the listener of the conversation, the party to the conversation Distribute the utterance data of the annotation candidates a request unit that requests annotation, and includes.

[0008] According to the first aspect, Data attribute in the determination unit, when performing annotation of a machine learning model that estimates the state of a party to the conversation from the utterance data of the conversation, Determine the attribute of the utterance data of the conversation from the utterance data, collate the determined attribute with pre-generated insufficient data to determine whether the utterance data corresponds to the insufficient data, and in the determination unit, use a model for pre-label determination to determine a provisional label for the utterance data corresponding to the insufficient data, and compare the determined provisional label with pre-determined insufficient label information to it is determined whether the utterance data is the utterance data of a predetermined annotation candidate that is insufficient.

[0009] Then, in the request unit, when it is determined by the determination unit that the utterance data is the annotation candidate utterance data that is insufficient, Based on the participant information of the conversation, determine the destination of the annotation request, at least one of the speaker and the listener of the conversation, the party to the conversation Distribute the utterance data of the annotation candidates is requested for annotation.

[0010] Thereby, it becomes possible to accurately perform annotation as compared with the case where a third party performs annotation of conversation data.

[0011] The information processing apparatus according to the second aspect is the information processing apparatus according to the first aspect, Further include an audio splitter that splits voice data containing a large number of utterances of meeting participants from the start to the end of the meeting into each utterance, and generates the data split into audio in units of utterances as the utterance data .

[0013] The information processing apparatus according to the third aspect Claim 1 or is the information processing apparatus according to the second aspect, The pre-generated insufficient data is data generated based on the target user of the machine learning model, the accuracy of the machine learning model, and the learning progress .

[0015] The information processing according to the fourth aspect Method isWhen a computer annotates a machine learning model for estimating the state of a party to a conversation from the utterance data of the conversation, it determines the attribute of the utterance data of the conversation from the utterance data, collates the determined attribute with pre-generated insufficient data to determine whether the utterance data corresponds to the insufficient data, a data attribute determination process; determines a provisional label of the utterance data corresponding to the insufficient data using a model for pre-label determination, compares the determined provisional label with pre-determined insufficient label information to determine whether the utterance data is the utterance data of an annotation candidate lacking annotation; and when it is determined by the determination process that the utterance data is the utterance data of the annotation candidate lacking annotation, determines a destination for the annotation request based on the participant information of the conversation, and distributes the utterance data of the annotation candidate to at least one of the speaker and the listener of the conversation, the parties to the conversation, to request an annotation, including a request process 。

[0016] According to the fourth aspect, It is possible to provide an information processing method capable of accurately performing annotation as compared with the case where a third party annotates conversation data 。

[0017] The information processing program according to the fifth aspect causes a computer to perform annotation of a machine learning model for estimating the state of a party to the conversation from the utterance data of the conversation When performing, it determines the attribute of the utterance data of the conversation from the utterance data, collates the determined attribute with pre-generated insufficient data to determine whether the utterance data corresponds to the insufficient data, a data attribute determination process; determines a provisional label of the utterance data corresponding to the insufficient data using a model for pre-label determination, and compares the determined provisional label with pre-determined insufficient label information to determine whether the utterance data is is insufficient is an to determine whether the utterance data is annotation candidate data a determination process to be performed When it is determined by the determination process that the utterance data is the annotation candidate data that is insufficient, Based on the participant information of the conversation, determine the destination for the annotation request and to perform a request process of requesting annotation from at least one of the speaker and the listener of the conversation, which are the parties to the conversation Distribute the utterance data of the annotation candidate and execute including the process .

[0018] According to the fifth aspect, it is possible to provide an information processing program capable of accurately performing annotation as compared with the case where a third party annotates conversation data.

Advantages of the Invention

[0019] As described above, according to the present invention, it is possible to provide an information processing apparatus and an information processing program capable of accurately performing annotation as compared with the case where a third party annotates conversation data. , information processing method and an information processing program.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Mode for Carrying Out the Invention

[0021] Hereinafter, an example of an embodiment of the present invention will be described in detail with reference to the drawings. FIG. 1 is a diagram showing a schematic configuration of the information processing system according to the present embodiment.

[0022] The information processing system 10 according to the present embodiment is a system that annotates a machine learning model for estimating the state of a conversation participant from conversation utterance data. Specifically, annotation of the machine learning model is performed by directly requesting annotation from the user who was present in the scene where the utterance was made.

[0023] As shown in FIG. 1, the information processing system 10 according to the present embodiment includes a recording device 12, a voice splitter 14, a DB (database) 16, a data attribute determiner 18, a pre-label determiner 20, an annotation distributor 22, and a UI (user interface) 24. Note that the data attribute determiner 18 and the pre-label determiner 20 correspond to an example of a determination unit, and the annotation distributor 22 corresponds to an example of a request unit.

[0024] The recording device 12 records or records the voice of a meeting or the like. The recording device 12 generates voice data by recording voice, for example, when a user uses a service or a web conferencing tool. Specifically, a recording function of a web conferencing tool or the like may be used.

[0025] Since the voice data of the recording device 12 contains a large number of utterances of meeting participants from the start to the end of the meeting, the voice splitter 14 splits the voice data containing a series of utterances into individual utterances and generates utterance data split into voice in units of utterances.

[0026] DB16 holds speech data, user information, meeting information, model information, annotation results, etc., and provides and stores the information necessary for various processes.

[0027] The data attribute determiner 18 determines the attributes of the speech data such as the speaker name for the segmented speech data.

[0028] The pre-label determiner 20 tentatively determines whether the target speech data contains missing labels.

[0029] The annotation distributor 22 distributes the speech data to the user for whom the annotation is requested and requests an annotation.

[0030] UI24 notifies the user of the speech data selected as the annotation target. After the user plays and checks the target speech data, the user selects an appropriate label on UI24 to perform the annotation. The annotation result and annotation information are stored in DB16 through UI24.

[0031] In this embodiment, the speech splitter 14, DB16, data attribute determiner 18, pre-label determiner 20, and annotation distributor 22 will be described as functions executed by a computer 11 (see FIG. 2), such as a server, which is an example of an information processing apparatus. FIG. 2 is a block diagram showing a schematic configuration of the computer.

[0032] As shown in FIG. 2, a computer 11 such as a server has a general computer configuration including a CPU (Central Processing Unit) 20A, a ROM (Read Only Memory) 20B, a RAM (Random Access Memory) 20C, a storage 20D, an interface (I / F) 20E, a bus 20F, and the like. Then, the CPU 11A expands and executes a program such as an information processing program stored in the ROM 11B in the RAM 11C, thereby functioning as a voice splitter 14, a data attribute determiner 18, a pre-label determiner 20, an annotation distributor 22, and a DB 16.

[0033] Subsequently, the data attribute determiner 18 will be described in detail. FIG. 3 is a block diagram showing the detailed configuration of the data attribute determiner 18.

[0034] As shown in FIG. 3, the data attribute determiner 18 has the functions of an attribute determination unit 28, a missing data management unit 30, a missing data determination unit 32, a priority determination unit 34, a data pool unit 36, and an output unit 38.

[0035] The attribute determination unit 28 determines the attribute of the input utterance data. For example, the gender of the utterance data may be determined as an attribute, or an individual may be identified by specifying the speaker name or the like. As an example, a method of identifying the speaker of the utterance data is used by a technique for identifying an individual such as voiceprint authentication.

[0036] The missing data management unit 30 holds the attributes of the learning data required for the machine learning model to be learned based on the information in the DB 16. Specifically, it holds the target user of the model (for whom the personal model is), the model accuracy, the learning progress status, etc., and generates missing data information based on this information.

[0037] The missing data determination unit 32 collates the attribute information determined by the attribute determination unit 28 with the missing data managed by the missing data management unit 30, and determines whether the utterance data corresponds to the missing data.

[0038] In the priority determination unit 34, when, as a result of the determination by the insufficient data determination unit 32, it corresponds to insufficient data, a priority comparison is performed with the annotation candidate data held in the data pool unit 36. And if the priority is higher than that of the held data, it is replaced with the data having a lower priority. Thereby, since the burden on the user becomes too large if the amount of annotation is too large, the upper limit amount of annotation can be controlled. Note that the priority is set to be higher as the degree of insufficiency of the insufficient data is higher.

[0039] The data pool unit 36 holds a predetermined number of annotation candidate data, and replaces the held annotation candidate data according to the determination by the priority determination unit 34.

[0040] In the output unit 38, for example, after the processing is completed for all the speech data during a meeting, the speech data to be annotated held in the data pool unit 36 is output.

[0041] Next, the pre-label discriminator 20 will be described in detail. FIG. 4 is a block diagram showing a detailed configuration of the pre-label discriminator 20.

[0042] As shown in FIG. 4, the pre-label discriminator 20 has functions of a label determination unit 40 as an example of a provisional label determination unit, an insufficient label management unit 42, an insufficient label determination unit 44, a priority determination unit 46, a data pool unit 48, and an output unit 50.

[0043] The pre-label discriminator 20 tentatively determines whether the speech data of the annotation candidate has a label with insufficient data.

[0044] In the label determination unit 40, using a model for preliminary label determination, the label of the input speech data of the annotation target is tentatively determined and output as temporary label information. For the label determination unit 40, the model with the highest accuracy among those capable of determining the label may be used. Here, since it is only necessary to know that the target speech data is likely to be the required label, a model with low accuracy or a model with a determination threshold adjusted differently from the actual application may be used.

[0045] Based on the information in the DB16, the insufficient label management unit 42 holds, as insufficient label information, the labels with low determination accuracy in the machine learning model to be learned.

[0046] The insufficient label determination unit 44 compares the insufficient label information obtained from the insufficient label management unit 42 with the temporary label information determined by the label determination unit 40, and determines whether the speech data of the annotation target corresponds to the insufficient label.

[0047] If, as a result of the determination by the insufficient label determination unit 44, it corresponds to an insufficient label, the priority determination unit 46 compares the priority with the speech data of the annotation target held in the data pool unit 48, and if the priority is higher than that of the held data, it exchanges it with the data with a lower priority. Note that the priority is set such that the higher the degree of insufficiency of the insufficient label, the higher the priority.

[0048] The data pool unit 48 holds a predetermined number of speech data of the annotation target, and exchanges the speech data of the annotation target to be held according to the determination of the priority determination unit 46.

[0049] After the processing is completed for all the speech data of the annotation target, the output unit 50 outputs the speech data of the annotation target held in the data pool unit 48.

[0050] Next, the annotation distributor 22 will be described in detail. FIG. 5 is a block diagram showing the detailed configuration of the annotation distributor 22.

[0051] The annotation distributor 22 has the functions of a distribution destination determination unit 52, a distribution content determination unit 54, and an output unit 56 as an example of a determination unit.

[0052] The distribution destination determination unit 52 refers to the meeting participant information in the DB 16 and determines the meeting participants who are the distribution destinations of the utterance data to be annotated. The distribution destination judges by considering whether the model under learning is a speaker model, a listener model, or both. The distribution to listeners may be for all or some of the meeting participants. When distributing to a part, adjustments such as preventing the annotation amount from being biased may be made.

[0053] The distribution content determination unit 54 determines the distribution content. For example, when distributing to the speaker himself / herself of the utterance and the listener who heard the utterance, the annotation request message is changed.

[0054] The output unit 56 outputs the distribution content determined by the distribution content determination unit 54 to the UI 24. As a specific distribution method, for example, a chat tool or a dedicated application is used.

[0055] The meeting participant who received the annotation request plays the utterance data on the UI 24 such as a chat tool and answers an appropriate label. For example, when wanting to annotate whether the utterance was positive, negative, or normal, on the UI 24 of the chat tool, a playback button for the utterance data and selection buttons for positive, negative, and normal corresponding labels are displayed. Then, the meeting participant operates the corresponding selection button to perform the annotation. The label obtained by the annotation is stored in the DB 16 as annotation information, associated with the utterance data and the annotation implementer.

[0056] Next, specific processing performed by a computer 11 such as a server of the information processing system 10 according to the present embodiment configured as described above will be described. FIG. 6 is a flowchart showing an example of the processing flow performed by the computer 11 of the information processing system 10 according to the present embodiment. Note that the processing in FIG. 6 starts when, for example, a user operates the computer 11 to acquire voice data from the recording device 12 and an annotation request start is instructed.

[0057] In step 100, the CPU 11A divides the voice data obtained from the recording device 12 for each utterance and proceeds to step 102. That is, the voice splitter divides the voice data including a series of utterances for each utterance and generates utterance data divided into voice in units of utterances.

[0058] In step 102, the CPU 11A focuses on one piece of divided utterance data, identifies the speaker of the utterance data, and proceeds to step 104. That is, the attribute determination unit 28 identifies the speaker by determining the attribute of the input utterance data.

[0059] In step 104, the CPU 11A determines whether the utterance data is insufficient-attribute utterance data. This determination is made by the insufficient data determination unit 32 comparing the attribute information determined by the attribute determination unit 28 with the insufficient data managed by the insufficient data management unit 30 to determine whether the utterance data corresponds to insufficient data. If this determination is affirmed, the process proceeds to step 106; if it is negated, the process proceeds to step 112.

[0060] In step 106, the CPU 11A determines the priority of the utterance data and proceeds to step 108. That is, the priority determination unit 34 determines the priority by comparing the priority with the annotation candidate data held in the data pool unit 36.

[0061] In step 108, the CPU 11A determines whether the priority is higher than that of the existing data. This determination is made by the priority determination unit 34 to determine whether the priority is higher than that of the annotation candidate data held in the data pool unit 36. If this determination is affirmed, the process proceeds to step 110; if it is negated, the process proceeds to step 112.

[0062] In step 110, the CPU 11A saves the utterance data in the data pool unit 36 as annotation candidate data and proceeds to step 112. That is, the priority determination unit 34 replaces the data with low priority among the annotation candidate data held in the data pool unit 36.

[0063] In step 112, the CPU 11A determines whether the processing of all data has been completed. This determination is made by the output unit 38 to determine whether the processing of steps 102 to 110 has been completed for all of the divided utterance data. If this determination is negated, the process returns to step 102 above and focuses on the next utterance data to perform the above processing. On the other hand, if the determination is affirmed, the process proceeds to step 114.

[0064] In step 114, the CPU 11A focuses on one piece of divided utterance data, determines the label of the utterance data, and proceeds to step 116. That is, the label determination unit 40 tentatively determines the label of the input utterance data to be annotated using a model for pre-label determination and outputs it as temporary label information.

[0065] In step 116, the CPU 11A determines whether the utterance data is insufficient-labeled data. This determination is made by the insufficient-label determination unit 44 by comparing the insufficient-label information obtained from the insufficient-label management unit 42 with the temporary label information determined by the label determination unit 40 to determine whether the utterance data is insufficient-labeled data. If this determination is affirmed, the process proceeds to step 118; if it is negated, the process proceeds to step 124.

[0066] In step 118, the CPU 11A determines the priority of the target utterance data and proceeds to step 120. That is, the priority determination unit 46 determines the priority by comparing the priority with the utterance data to be annotated held in the data pool unit 48.

[0067] In step 120, the CPU 11A determines whether the priority is higher than that of the existing data. This determination is made by the priority determination unit 46 to determine whether the priority is higher than that of the utterance data to be annotated held in the data pool unit 48. If this determination is affirmed, the process proceeds to step 122; if it is negated, the process proceeds to step 124.

[0068] In step 122, the CPU 11A saves the utterance data in the data pool unit 48 as the utterance data to be annotated and proceeds to step 112. That is, the priority determination unit 46 replaces the data with low priority among the utterance data to be annotated held in the data pool unit 48.

[0069] In step 124, the CPU 11A determines whether the processing of all data has been completed. This determination is made by the output unit 50 to determine whether the processing in steps 114 to 122 has been completed for all of the divided utterance data. If this determination is negated, the process returns to step 114 above and focuses on the next utterance data to perform the above processing. On the other hand, if the determination is affirmed, the process proceeds to step 126.

[0070] In step 126, the CPU 11A focuses on one of the divided utterance data, determines the distribution destination of the utterance data, and proceeds to step 128. That is, the distribution destination determination unit 52 refers to the conference participant information in the DB 16 and determines the conference participant who is the distribution destination of the utterance data to be annotated.

[0071] In step 128, the CPU 11A determines whether the destination is the speaker himself / herself. This determination is made by the destination determination unit 52 to determine whether the determined destination is the speaker himself / herself. If this determination is affirmed, the process proceeds to step 130; if it is negated, the process proceeds to step 132.

[0072] In step 130, the CPU 11A generates a delivery message for the speaker and proceeds to step 134. That is, the delivery content determination unit 54 generates a delivery message for the speaker as the delivery content.

[0073] On the other hand, in step 132, the CPU 11A generates a delivery message for the listener and proceeds to step 134. That is, the delivery content determination unit 54 generates a delivery message for the listener as the delivery content.

[0074] In step 134, the CPU 11A determines whether the processing of all data is completed. This determination is made by the output unit 56 to determine whether the processing in steps 126 to 132 has been completed for all of the divided utterance data. If this determination is negated, the process returns to step 126 above and focuses on the next utterance data to perform the above processing. On the other hand, if the determination is affirmed, the process proceeds to step 136.

[0075] In step 136, the CPU 11A distributes an annotation request to end a series of processes. That is, the output unit 56 outputs the delivery content determined by the delivery content determination unit 54 to the UI 24. Then, the conference participant who has received the annotation request plays back the utterance data on the UI 24 such as a chat tool and answers with an appropriate label.

[0076] As described above, in the computer 11 of the information processing system 10 according to the present embodiment, when annotating a machine learning model that estimates the state of a conversation participant from the utterance data of a conversation, a determination process for determining whether the utterance data is the predetermined annotation candidate data that is lacking is performed, and when it is determined by the determination process that the utterance data is the annotation candidate data that is lacking, a request process for requesting an annotation from at least one of the speaker and the listener of the conversation, i.e., the conversation participant, is performed. Thereby, compared with the case where a third party annotates the conversation data, the annotation can be accurately implemented.

[0077] In addition, by determining the provisional label of the utterance data of the conversation and determining whether the utterance data is the annotation candidate data that is lacking by using the provisional label, it becomes possible to request an annotation only for the necessary data, and the burden on the user can be reduced.

[0078] In addition, by performing an attribute determination for determining the attribute of the utterance data of the conversation and determining whether the utterance data is the annotation candidate data that is lacking by using the attribute and the provisional label, it becomes possible to further narrow down and request the data for which an annotation is requested, and the burden on the user can be further reduced.

[0079] Furthermore, based on the participant information of the conversation, the destination of the annotation request is determined and the annotation is requested, so that it becomes possible to surely request the annotation to the conversation participant.

[0080] In the above-described embodiments, the voice splitter 14, the DB 16, the data attribute determiner 18, the pre-label determiner 20, and the annotation distributor 22 have been described as functions executed by a single computer 11 such as a server. However, the present invention is not limited to this. For example, it may be configured by a computer such as a server having the functions of the voice splitter 14, the data attribute determiner 18, the pre-label determiner 20, and the annotation distributor 22, and a server having the function of the DB 16. Alternatively, each function of the voice splitter 14, the DB 16, the data attribute determiner 18, the pre-label determiner 20, and the annotation distributor 22 may be executed by separate computers, or each function may be provided by sharing among three or more computers and executed in cooperation between the computers.

[0081] Also, the processing performed by the computer 11 of the information processing system 10 in each of the above-described embodiments has been described as software processing performed by executing a program. However, the present invention is not limited to this. For example, the processing may be performed by hardware such as a GPU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), and an FPGA (Field-Programmable Gate Array). Alternatively, the processing may be a combination of both software and hardware. Further, when the processing is software processing, the program may be stored in various storage media and distributed.

[0082] Furthermore, the present invention is not limited to the above, and it goes without saying that various modifications can be made and implemented within the scope without departing from the gist of the present invention.

Description of Reference Numerals

[0083] 10 Information processing system 11 Computer 12 Recording device 14 Voice splitter 16 DB 18 Data attribute determiner (determination unit) 20 Pre - label Determinator (Determination Unit) 22 Annotation Distributor (Requesting Unit) 24 UI 28 Attribute Determination Unit 30 Insufficient Data Management Unit 32 Insufficient Data Determination Unit 40 Label Determinator (Provisional Label Determinator) 42 Insufficient Label Management Unit 44 Insufficient Label Determination Unit 52 Destination Determination Unit 54 Distribution Content Determination Unit

Claims

1. When performing annotation of a machine learning model for estimating the state of a party to the conversation from the utterance data of the conversation, determining an attribute of the utterance data of the conversation from the utterance data, and collating the determined attribute with pre-generated insufficient data to determine whether the utterance data corresponds to the insufficient data, a data attribute determination unit; Determining a provisional label of the utterance data corresponding to the insufficient data using a model for pre-label determination, and comparing the determined provisional label with pre-determined insufficient label information to determine whether the utterance data is utterance data of an annotation candidate lacking annotation; a determination unit; When it is determined by the determination unit that the utterance data of the annotation candidate is lacking, determining a destination for annotation request based on the participant information of the conversation, and distributing the utterance data of the annotation candidate to at least one of the speaker and the listener of the conversation, the parties to the conversation, and requesting annotation; a request unit; An information processing apparatus comprising the above.

2. The information processing apparatus according to claim 1, further comprising an audio splitter that splits audio data including a large number of utterances of conference participants from the start to the end of the conference into individual utterances, and generates data split into utterance-level audio as the utterance data.

3. The information processing apparatus according to claim 1, wherein the pre-generated insufficient data is data generated based on the target user of the machine learning model, the accuracy of the machine learning model, and the learning progress status.

4. A computer performs When performing annotation of a machine learning model for estimating the state of a party to the conversation from the utterance data of the conversation, determining an attribute of the utterance data of the conversation from the utterance data, and collating the determined attribute with pre-generated insufficient data to determine whether the utterance data corresponds to the insufficient data, a data attribute determination process; Determine a provisional label for the utterance data corresponding to the missing data using a model for pre-label determination, compare the determined provisional label with pre-determined missing label information, and perform a determination process to determine whether the utterance data is the utterance data of an annotation candidate lacking the annotation; When it is determined by the determination process that the utterance data is the utterance data of the annotation candidate lacking the annotation, based on the participant information of the conversation, determine the destination of the annotation request, and distribute the utterance data of the annotation candidate to at least one of the speaker and the listener of the conversation, the parties to the conversation, and perform a request process to request an annotation; An information processing method for performing a process including the above.

5. On a computer, When annotating a machine learning model that estimates the state of the parties to the conversation from the utterance data of the conversation, determine the attribute of the utterance data of the conversation from the utterance data, and collate the determined attribute with pre-generated missing data to determine whether the utterance data corresponds to the missing data, a data attribute determination process; Determine a provisional label for the utterance data corresponding to the missing data using a model for pre-label determination, compare the determined provisional label with pre-determined missing label information, and perform a determination process to determine whether the utterance data is the utterance data of an annotation candidate lacking the annotation; When it is determined by the determination process that the utterance data is the utterance data of the annotation candidate lacking the annotation, based on the participant information of the conversation, determine the destination of the annotation request, and distribute the utterance data of the annotation candidate to at least one of the speaker and the listener of the conversation, the parties to the conversation, and perform a request process to request an annotation; An information processing program for causing the above processes to be executed.

Citation Information

Patent Citations

  • Information processor, information processing method, and program

    JP2018106662A

  • JPP7152825B

  • Learning device, space control device, learning program, and space control program

    WO2020246600A1