Information processing device, information processing method, and information processing program

JPWO2025126418A5Pending Publication Date: 2026-09-01
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025563171
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2026-06-04
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

Existing techniques for detecting artificially generated voices using deep learning suffer from low detection accuracy when inputs belong to a domain different from the training data, and fine-tuning requires a large amount of training data and is costly.

Method used

An information processing apparatus and method that processes training data using a prompt and updates the prompt to minimize the difference between predicted values from a detector and the attached labels, thereby improving detection accuracy without needing extensive training data.

Benefits of technology

The approach enhances the detection accuracy of synthetic speech using a detector without requiring a large amount of training data, addressing the limitations of existing methods.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This information processing device comprises: a processing unit that uses a prompt to process training data which represents speech and to which a label is attached; and an updating unit that updates the prompt so as to reduce a difference between a predicted value obtained by inputting the training data processed by the processing unit to a detector for detecting artificial speech, and the label attached to the training data.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and information processing program

[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing program.

[0002] Techniques for detecting artificially generated voices (hereinafter also referred to as "artificial voices") using deep learning techniques have been proposed. For example, Patent Literature 1 discloses a system and method that implements a neural network architecture for detecting spoofing in audio signals.

[0003] Special Publication No. 2023-511104

[0004] However, the technology described in Patent Document 1 has a problem in that, for example, detection accuracy is low for inputs belonging to domains (language type, sound collection environment, model generation method, etc.) different from the training data used in training the neural network. Although fine tuning is an example of a method for improving detection accuracy, fine tuning has the problem of requiring a large amount of training data and being expensive.

[0005] The present disclosure has been made in consideration of the above-mentioned problems, and an exemplary purpose thereof is to provide a technology for improving the detection accuracy of an artificial voice using a detector that detects an artificial voice without requiring a large amount of training data.

[0006] An information processing device according to an exemplary aspect of the present disclosure includes: a processing means for processing training data representing labeled speech using a prompt; and an update means for updating the prompt so as to reduce the difference between a predicted value obtained by inputting the training data processed by the processing means into a detector that detects artificial speech and the label attached to the training data.

[0007] An information processing method according to an exemplary aspect of the present disclosure includes: a processing step in which at least one processor processes training data representing labeled speech using a prompt; and an update step in which the at least one processor updates the prompt so as to reduce a difference between a predicted value obtained by inputting the training data processed in the processing step into a detector that detects artificial speech and the label assigned to the training data.

[0008] An information processing program according to an exemplary aspect of the present disclosure is an information processing program that causes a computer to function as an information processing device, and causes the computer to function as: a processing means that processes training data representing labeled speech using prompts; and an update means that updates the prompts so as to reduce the difference between a predicted value obtained by inputting the training data processed by the processing means into a detector that detects artificial speech and the label assigned to the training data.

[0009] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that a technology can be provided that improves the detection accuracy of an artificial voice using a detector that detects an artificial voice without requiring a large amount of training data.

[0010] FIG. 1 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 2 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 3 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 4 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 5 is a diagram showing an overview of prompt tuning according to the present disclosure. FIG. 6 is a diagram showing an overview of audio deepfake detection according to the present disclosure. FIG. 7 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 8 is a diagram showing an outline of an example of processing processing according to the present disclosure. FIG. 9 is a flow diagram showing an example of the flow of a learning method according to the present disclosure. FIG. 10 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 11 is a flow diagram showing an example of the flow of a discrimination method according to the present disclosure. FIG. 12 is a block diagram showing the configuration of a computer functioning as an information processing device according to the present disclosure.

[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0012] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0013] (Configuration of Information Processing Apparatus) The configuration of the information processing apparatus 1 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing apparatus 1. As shown in Fig. 1, the information processing apparatus 1 includes a processing unit 11 and an updating unit 12.

[0014] The processing unit 11 processes training data representing labeled speech using prompts. The update unit 12 updates the prompts so as to reduce the difference between a predicted value obtained by inputting the training data processed by the processing unit 11 into a detector that detects artificial speech and the label attached to the training data.

[0015] (Effects of the Information Processing Device) As described above, the information processing device 1 employs a configuration including a processing unit 11 that processes training data representing labeled speech using prompts, and an update unit 12 that updates the prompts so as to reduce the difference between a predicted value obtained by inputting the training data processed by the processing unit 11 into a detector that detects artificial speech and the label attached to the training data. Therefore, the information processing device 1 has the effect of improving the detection accuracy of artificial speech using a detector that detects artificial speech without requiring a large amount of training data.

[0016] (Flow of Information Processing Method) The flow of the information processing method S1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the information processing method S1. As shown in Fig. 2, the information processing method S1 includes a processing process S11 and an update process S12.

[0017] In a processing step S11, at least one processor processes training data representing labeled speech using a prompt. In an updating step S12, at least one processor updates the prompt so as to reduce the difference between a predicted value obtained by inputting the training data processed in the processing step S11 into a detector for detecting artificial speech and the label assigned to the training data.

[0018] (Effects of the Information Processing Method) As described above, the information processing method S1 employs a configuration including: a processing step in which at least one processor processes training data representing labeled speech using a prompt; and an update step in which the at least one processor updates the prompt so as to reduce the difference between a predicted value obtained by inputting the training data processed in the processing step into a detector that detects artificial speech and the label assigned to the training data. Therefore, the information processing method S1 has the effect of improving the detection accuracy of artificial speech using a detector that detects artificial speech without requiring a large amount of training data.

[0019] (Configuration of Information Processing Device) The configuration of the information processing device 2 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of the information processing device 2. As shown in Fig. 3, the information processing device 2 includes a processing unit 21 and a discrimination unit 22.

[0020] The processing unit 21 processes the voice data to be detected using a prompt. The prompt is a prompt that has been updated so as to reduce the difference between (i) a predicted value obtained by inputting training data representing labeled voice and processed using the prompt into a detector that detects artificial voices, and (ii) the label assigned to the training data. The determination unit 22 inputs the voice data processed by the processing unit 21 into the detector, thereby determining whether the voice represented by the voice data to be detected is an artificial voice.

[0021] (Effects of the Information Processing Device) As described above, the information processing device 2 employs a configuration including: (i) a processing unit 21 that processes speech data to be detected using a prompt that has been updated so as to reduce the difference between a predicted value obtained by inputting training data representing labeled speech and processed using a prompt into a detector that detects artificial speech, and (ii) the label assigned to the training data; and a discrimination unit 22 that inputs the speech data processed by the processing unit 21 into the detector and determines whether the speech represented by the speech data to be detected is an artificial speech. Therefore, the information processing device 2 provides the effect of improving the detection accuracy of artificial speech using a detector that detects artificial speech without requiring a large amount of training data.

[0022] (Flow of Information Processing Method) The flow of the information processing method S2 will be described with reference to Fig. 4. Fig. 4 is a flow chart showing the flow of the information processing method S2. As shown in Fig. 4, the information processing method S2 includes a processing process S21 and a determination process S22.

[0023] In the processing step S21, at least one processor processes the target speech data using a prompt that has been updated so as to reduce the difference between (i) a predicted value obtained by inputting training data representing labeled speech, which has been processed using a prompt, into a detector that detects artificial speech, and (ii) the label assigned to the training data. In the determination step S22, at least one processor inputs the speech data processed in the processing step S21 into the detector, thereby determining whether the speech represented by the target speech data is an artificial speech.

[0024] (Effects of Information Processing Method) As described above, the information processing method S2 employs a configuration including: (i) a processing process S21 in which at least one processor processes speech data to be detected using a prompt that has been updated so as to reduce the difference between (i) a predicted value obtained by inputting training data representing labeled speech and that has been processed using a prompt into a detector that detects artificial speech, and (ii) a label assigned to the training data; and (ii) a determination process S22 in which the at least one processor inputs the speech data processed in the processing process S21 into the detector and determines whether the speech represented by the speech data to be detected is an artificial speech. Therefore, the information processing method S2 has the effect of improving the detection accuracy of artificial speech using a detector that detects artificial speech without requiring a large amount of training data.

[0025] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0026] (Overview of Information Processing Device) The information processing device 1A according to the present disclosure has a function of updating prompts used to detect artificial voices such as voice deepfakes. Voice deepfakes are elaborate artificial voices generated using deep learning technology. Voice deepfakes can be generated in real time and may be misused to impersonate others. Detectors for detecting artificial voices such as voice deepfakes include, but are not limited to, machine learning models that determine whether an input voice is "Fake" or "Real."

[0027] Prompts are data used to process the audio data that is input to the detector. Learning prompts using training data after fixing the detector's model parameters is also called "prompt tuning."

[0028] 5 is a diagram showing an overview of prompt tuning. In FIG. 5, the information processing device 1A performs a processing process S100 in which training data x is processed using a prompt P. The information processing device 1A also updates the prompt P so as to improve detection accuracy based on the results obtained by inputting the data obtained by the processing process S100 into a detector 3 that detects artificial voices.

[0029] The training dataset D used for prompt tuning may include training data belonging to a domain different from the domain of the training data used for machine learning of the detector 3. Here, the term "domain" refers to the distribution of data. In the case of speech data, differences in domain arise due to, for example, differences in language, recording environment, and type of generation method, but are not limited to these. In the following description, the domain of the training data used for learning the detector 3 is also referred to as the "source domain." Furthermore, the domain of the training data used for prompt tuning is also referred to as the "target domain." The source domain is an example of a first domain according to the present disclosure, and the target domain is an example of a second domain according to the present disclosure.

[0030] As an example, the training data used in prompt tuning may be audio data in a language (e.g., Chinese) different from the language (e.g., English) of the training data used to train the detector 3. Furthermore, the training data may be audio data collected in a recording environment different from the recording environment of the audio represented by the training data used to train the detector 3.

[0031] 6 is a diagram illustrating an overview of artificial voice detection using prompts learned by the information processing device 1A. In the example of FIG. 6, three prompts PA, PB, and PC are used. The prompts PA, PB, and PC each correspond to a different domain, with prompt PA corresponding to domain DA, prompt PB corresponding to domain DB, and prompt PC corresponding to domain DC. Among the prompts PA, PB, and PC, a prompt corresponding to the domain of the voice data TD to be detected is selected and used. The domain of the voice data TD is classified, for example, by a domain classification model that classifies domains.

[0032] As an example, if the domain of the voice data TD is the domain DB, a prompt PB in the domain DB is selected from the prompts PA, PB, and PC, and a processing process S200 is performed to process the voice data TD using the selected prompt PB. The data obtained by the processing process S200 is input to the detector 3, which determines whether the voice data TD is an artificial voice. On the other hand, if the domain of the voice data TD is the source domain, processing using the prompts PA, PB, or PC is not performed. In this case, the voice data TD is input to the detector 3 without being processed.

[0033] (Configuration of information processing device) The configuration of the information processing device 1A will be described with reference to Fig. 7. Fig. 7 is a block diagram showing the configuration of the information processing device 1A. The information processing device 1A includes a control unit 10A, a storage unit 20A, a communication unit 30A, an input unit 40A, and an output unit 50A.

[0034] (Communication Unit) The communication unit 30A communicates with devices external to the information processing device 1A via a communication line. While the specific configuration of the communication line does not limit the present exemplary embodiment, examples of the communication line include a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination thereof. The communication unit 30A transmits data supplied from the control unit 10A to other devices, and supplies data received from other devices to the control unit 10A.

[0035] (Input Unit) The input unit 40A is configured to receive input to the information processing device 1A, and includes, for example, input devices such as a keyboard, a mouse, a touch panel, a camera, a microphone, etc. The input unit 40A may also be configured to receive data from the input devices via an interface such as a USB (Universal Serial Bus).

[0036] (Output Unit) The output unit 50A is a component for performing output from the information processing device 1A, and includes, for example, output devices such as a display, a printer, a touch panel, a speaker, etc. The output unit 50A may be configured to include, for example, an interface such as a USB, and to output data to the output device via the interface.

[0037] (Storage Unit) The storage unit 20A stores various types of information referenced by the control unit 10A. Examples of such information include prompts P1, P2, ..., Pn (n is a natural number equal to or greater than 1), a training data set D, and a detector parameter θ. Hereinafter, when there is no need to distinguish between the prompts P1, P2, ..., Pn, for convenience of explanation, they will be referred to as "prompt p."

[0038] (Prompt) The prompt p is data used to process the training data x to be input to the detector 3. As an example, the prompt p is k d-dimensional vectors, but is not limited to this.

[0039] (Training Data Set) The training data set D is a set of training data x labeled with y. The training data x is training data used for prompt tuning. As an example, the training data set D is expressed by the following formula (1). In equation (1), x is voice data, and y∈{Fake, Real} is a label. The label indicates whether the voice represented by the training data x is an artificial voice. "Fake" indicates that the voice represented by the training data x is an artificial voice, and "Real" indicates that the voice represented by the training data x is not an artificial voice. M is the total number of training data x.

[0040] As an example, the training data set D includes training data that belongs to a domain different from a source domain, which is the domain of the training data used to train the detector 3. The training data set D may also include training data that belongs to the source domain.

[0041] (Detector Parameter) The detector parameter θ is a model parameter that defines the detector 3. The detector 3 is a detector that detects artificial speech. As an example, the detector 3 is a model generated by machine learning using training data belonging to the source domain, but is not limited to this. An example of the detector 3 is a Transformer.

[0042] (Control Unit) The control unit 10A includes a processing unit 11A, a loss calculation unit 12A, a prompt update unit 13A, and a detector parameter update unit 14A. The processing unit 11A is an example of a processing means according to the present disclosure. The loss calculation unit 12A, the prompt update unit 13A, and the detector parameter update unit 14A are examples of an update means according to the present disclosure.

[0043] (Processing Unit) The processing unit 11A acquires training data x and a prompt p and processes the acquired training data x using the prompt p. As an example, the processing unit 11A may acquire the training data x and / or the prompt p by reading the training data x and / or the prompt p from a storage destination (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A) designated by the user of the information processing device 1A. Alternatively, the processing unit 11A may acquire the training data x and / or the prompt p by receiving the training data x and / or the prompt p from another device via the communication unit 30A. Alternatively, the processing unit 11A may acquire the training data x and / or the prompt p input to the input unit 40A.

[0044] When the domain of the training data x is different from the source domain, it can also be said that the processing unit 11A processes, using the prompt p, the training data x that belongs to a domain (second domain) different from the domain (first domain) of the training data used to train the detector 3. Furthermore, when the training data D includes both training data of the source domain and training data of the target domain, it can also be said that the processing unit 11A processes the training data that belongs to the source domain in addition to the training data that belongs to the target domain.

[0045] Furthermore, if there are multiple prompts p, the processing unit 11A selects a prompt p that corresponds to the domain of the training data x from the multiple prompts p, and processes the training data x using the selected prompt p.

[0046] As an example, the processing unit 11A performs at least one of a process of adding a prompt p to the training data x and a process of combining the training data x. However, the processing processes performed by the processing unit 11A are not limited to these. In the case of combining, as an example, the processing unit 11A performs a process of adding a prompt p to the beginning of the training data x. However, the position where the processing unit 11A adds the prompt p may be a position other than the beginning of the training data x.

[0047] (Loss Calculation Unit) The loss calculation unit 12A calculates a loss L that represents the difference between a predicted value obtained by inputting the training data x processed by the processing unit 11A to the detector 3 and a label y assigned to the training data x. An example of the loss L is cross-entropy loss, but is not limited to this.

[0048] As an example, the loss calculation unit 12A calculates the loss L using the following formula: L=Loss(x, y x ) where Loss() is the loss function, x is the training data, y x is the label attached to the training data x. In this case, the training data set D is the training data x belonging to the target domain. t and the training data x belonging to the source domain s If the training data x s The predicted value (first predicted value) obtained by s The label y attached to xs and the first difference between the training data x t The predicted value (second predicted value) obtained by t The label y attached to xt and the second difference are calculated using a common loss function.

[0049] As another example, the loss calculation unit 12A may calculate the loss L using the following formula: L=Loss(x s , yxs ) + αLoss(x t , y xt ) where α is a weight hyperparameter. In this case, the loss calculation unit 12A and the prompt update unit 13A calculate the training data x s The first predicted value obtained by s The label y attached to xs and the first difference between the training data x t The second predicted value obtained by t The label y attached to xt It can also be said that the prompt p is updated using a weighted sum of the first difference and the second difference.

[0050] (Prompt Updater) The prompt updater 13A updates the prompt p so as to reduce the loss L. For example, if the prompt p is a vector, the prompt updater 13A adjusts the component values ​​of the vector so as to reduce the loss L, as an example.

[0051] The prompt updating unit 13A outputs the updated prompt. For example, the prompt updating unit 13A may output the prompt p by writing the updated prompt p to a storage destination (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A) designated by the user of the information processing device 1A. The prompt updating unit 13A may also output the prompt p by transmitting the prompt p via the communication unit 30A, or may output the prompt p to an output device such as a display.

[0052] (Detector Parameter Updater) The detector parameter updater 14A updates at least some of the detector parameters θ so as to reduce the difference between the predicted value and the label. For example, the detector parameter updater 14A may update only the model parameters of the final layer of the model (e.g., a linear layer of a transformer). The loss function used to update the model parameters may be the same as the loss function used to update the prompt p, or may be a different loss function.

[0053] Furthermore, the detector parameter update unit 14A may update the additional model parameters of the detector 3 so as to reduce the difference between the predicted value and the label. Here, the additional model parameters are parameters that are added to the detector 3, such as parameters added to a Transformer by LoRA (Low-Rank-Adaptation). LoRA is a method for performing additional learning with a small amount of computation, in which a difference matrix is ​​added next to a linear layer (such as the attention mechanism of a Transformer) and is used as an update target. The loss function used to update the additional model parameters may be the same as the loss function used to update the prompt p, or may be a different loss function.

[0054] The detector parameter update unit 14A may output the updated parameters. As an example, the detector parameter update unit 14A may output the parameters by writing the updated parameters to a storage destination (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A) designated by the user of the information processing device 1A. Furthermore, the detector parameter update unit 14A may output the parameters by transmitting the parameters via the communication unit 30A, or may output the parameters to an output device such as a display.

[0055] (Specific Example of Processing) Fig. 8 is a diagram schematically illustrating a specific example of processing performed by the processing unit 11A. In the example of Fig. 8, the detector 3 is a Transformer, and includes a Transformer Encoder 31, a Head 32, and a Linear Layer 33. The input to the detector 3 is a token, and the processing unit 11A performs at least one of a process of adding a prompt Pa to the token obtained from the training data x and a process of combining a prompt Pb.

[0056] 8 , the processing unit 11A first generates tokens T representing the feature quantities of each of the multiple divisions of training data x. As an example, the tokens T are vectors representing the feature quantities. As an example, the processing unit 11A divides the training data x, which is speech data representing speech, into k pieces (k is a natural number of 2 or more), and inputs the multiple pieces of data obtained by the division into a feature extractor 4 that extracts feature quantities, thereby generating k d-dimensional feature quantities representing the speech features as tokens T.

[0057] The feature extractor 4 is a model that outputs feature quantities of input data, and is, for example, a trained model generated by machine learning. The feature extractor 4 may be stored in the storage unit 20A of the information processing device 1A, or may be stored in a storage device other than the information processing device 1A. When the feature extractor 4 is stored in a storage device other than the information processing device 1A, the processing unit 11A may, for example, input data obtained by dividing the training data x to the feature extractor 4 by transmitting the data via the communication unit 30A or outputting the data from the output unit 50A. Note that, "the feature extractor 4 is stored in a storage device" means that parameters that define the feature extractor 4 are stored in the storage device.

[0058] Next, the processing unit 11A processes tokens T, which represent each of the feature quantities obtained by dividing the training data x into multiple parts, using prompts Pa and / or Pb. When performing an addition process to add prompt Pa to token T, the processing unit 11A generates k d-dimensional feature quantities by adding prompt Pa, which is k d-dimensional feature quantities, to token T, which is also k d-dimensional feature quantities. When performing a combination process to combine prompt Pb with token T, the processing unit 11A generates (k+N) d-dimensional feature quantities by combining prompt Pb, which is N d-dimensional feature quantities (N is a natural number greater than or equal to 1), with token T, which is also k d-dimensional feature quantities.

[0059] The processing unit 11A inputs the token T processed using the prompt Pa and / or Pb to the detector 3. The detector 3 outputs a prediction value indicating whether the input token T is an artificial voice.

[0060] (Flow of Learning Method) FIG. 9 is a flow diagram showing an example of the flow of a learning method S1A executed by the information processing device 1A. Some of the steps included in the flow diagram of FIG. 9 may be executed in parallel or in a different order. In step S101, the processing unit 11A acquires training data x. In addition, in step S102, the processing unit 11A acquires a detector 3. Note that acquiring the detector 3 means acquiring a detector parameter θ. In step S103, the processing unit 11A initializes a prompt p. As an example, the processing unit 11A initializes the prompt p using a random number.

[0061] In step S104, the processing unit 11A processes training data x representing speech with a label y using a prompt p. In S105, the loss calculation unit 12A calculates a loss L using a predicted value obtained by inputting the training data x processed in step S104 to the detector 3 and the label y assigned to the training data x. In S106, the prompt update unit 13A updates the prompt p so that the loss L becomes smaller.

[0062] In S107, the detector parameter update unit 14A updates at least some of the detector parameters θ of the detector 3 and / or the additional model parameters. In step S108, the detector parameter update unit 14A determines whether to end the learning process. As an example, the detector parameter update unit 14A may determine to end the learning process when the number of times the update process has been executed reaches a predetermined number. If the learning process is not to be ended (NO in step S108), the detector parameter update unit 14A returns to the process of step S104. If the learning process is to be ended (YES in step S108), the detector parameter update unit 14A ends the process.

[0063] (Effects of Information Processing Device) However, there is a problem with artificial speech detectors in that their detection performance deteriorates for inputs belonging to domains different from the source domain. For example, detection accuracy decreases for fake data generated in a language different from that of the training data used to train the detector, or generated using a different process. Fine-tuning (additional learning) is one method for improving detection accuracy for the target domain, but fine-tuning requires a large amount of training data and has the disadvantage of being expensive. Furthermore, fine-tuning can lead to a deterioration in prediction accuracy for the already trained source domain, which is known as destructive forgetting.

[0064] In the information processing device 1A, the detector 3 is a model generated by machine learning using training data belonging to a source domain, and the processing unit 11A is configured to process training data belonging to a target domain different from the source domain using a prompt p. This makes it possible to improve the detection accuracy for input data belonging to the target domain.

[0065] Furthermore, the information processing device 1A employs a configuration in which the processing unit 11A executes at least one of a process of adding a prompt p to training data x and a process of combining the training data x. Therefore, according to the information processing device 1A, by learning the prompt p using a predicted value obtained from the training data x to which the prompt p has been added and / or combined, it is possible to improve the detection accuracy of the artificial voice using the detector 3 without requiring a large amount of training data.

[0066] Furthermore, the information processing device 1A employs a configuration in which the processing unit 11A generates tokens representing feature quantities from each of a plurality of parts obtained by dividing the training data x, and processes the generated tokens using the prompt p. Therefore, according to the information processing device 1A, by learning the prompt p using a predicted value obtained from the tokens processed using the prompt p, it is possible to improve the detection accuracy of the artificial voice using the detector 3 without requiring a large amount of training data.

[0067] Furthermore, in the information processing device 1A, the detector parameter update unit 14A is configured to update the detector parameter θ so that the difference between the predicted value and the label becomes smaller. Therefore, according to the information processing device 1A, it is possible to improve the detection accuracy of the detector 3 for inputs in the target domain.

[0068] Furthermore, in the information processing device 1A, the detector parameter update unit 14A is configured to update the additional model parameters of the detector 3 so as to reduce the difference between the predicted value and the label. Therefore, according to the information processing device 1A, it is possible to improve the detection accuracy of the detector 3 for inputs in the target domain.

[0069] Furthermore, the information processing device 1A employs a configuration in which a prompt p corresponding to a domain of training data x is selected from among a plurality of prompts p, and the training data x is processed using the selected prompt p. Therefore, the information processing device 1A can improve detection accuracy for inputs from a plurality of domains. In particular, the prompt p has the advantage of being small in data size, and compared to generating detectors 3 corresponding to each of the plurality of domains, it is possible to improve detection accuracy for inputs from a plurality of domains with a small data size.

[0070] In the information processing device 1A, the processing unit 11A processes the training data x belonging to the target domain. t In addition, the training data x belonging to the source domain s Therefore, the information processing device 1A can improve the detection accuracy for the input of the target domain while preventing a decrease in the detection accuracy for the input of the source domain.

[0071] In the information processing device 1A, the loss calculation unit 12A calculates the training data x belonging to the source domain. s The first predicted value obtained by s The label y attached to xs and the training data x belonging to the target domain. tThe second predicted value obtained by xt The information processing device 1A is configured to calculate the first difference and the second difference using a common loss function. Therefore, the information processing device 1A can increase the detection accuracy for the target domain input while preventing a decrease in the detection accuracy for the source domain input.

[0072] In the information processing device 1A, the loss calculation unit 12A calculates the training data x belonging to the source domain. s The first predicted value obtained by s The label y attached to xs and the training data x belonging to the target domain. t The second predicted value obtained by t The label y attached to xt The information processing device 1A is configured to update the prompt p using a weighted sum of the first difference and the second difference between ...

[0073] [Third Exemplary Embodiment] A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0074] (Configuration of Information Processing Device) The configuration of information processing device 1B will be described with reference to Fig. 10. Fig. 10 is a block diagram showing the configuration of information processing device 1B. Information processing device 1B switches between M (M is a natural number equal to or greater than 2) pre-created prompts p depending on the domain of the voice data.

[0075] The information processing device 1B includes a control unit 10B, a storage unit 20A, a communication unit 30A, an input unit 40A, and an output unit 50A. The control unit 10B includes a learning phase execution unit 111B and an estimation phase execution unit 112B. The learning phase execution unit 111B updates a plurality of prompts p. The learning phase execution unit 111B includes a processing unit 11A, a loss calculation unit 12A, a prompt update unit 13A, and a detector parameter update unit 14A.

[0076] The estimation phase execution unit 112B executes a process to determine whether the voice represented by the voice data is an artificial voice. The estimation phase execution unit 112B includes a domain determination unit 15B, a processing unit 16B, and a prediction unit 17B. The domain determination unit 15B is an example of a selection means according to the present disclosure. The processing unit 16B is an example of a second processing means according to the present disclosure. The prediction unit 17B is an example of a determination means according to the present disclosure.

[0077] (Domain Determination Unit) The domain determination unit 15B acquires speech data to be detected and selects a prompt p corresponding to the domain of the speech data from among multiple prompts p. The multiple prompts p are prompts updated by the prompt update unit 13A, and each prompt corresponds to a different domain. In other words, the prompt p is a prompt that has been updated so as to reduce the difference between the label attached to the training data and a predicted value obtained by inputting training data representing labeled speech and processed using the prompt into the detector 3 that detects artificial speech.

[0078] For example, the domain determination unit 15B acquires the voice data by reading the voice data from a storage location (which may be a storage device within the information processing device 1B or a storage device external to the information processing device 1B) designated by the user of the information processing device 1B. Alternatively, the domain determination unit 15B may acquire the voice data by receiving the voice data from another device via the communication unit 30A. Alternatively, the domain determination unit 15B may acquire the voice data input to the input unit 40A.

[0079] An example of a method for determining the domain of voice data is a method that uses the prediction results of a machine learning model that performs domain classification. In this case, an example of the machine learning model is a model (a model that performs (M+1) class classification) that classifies the domain (source domain) of the training data of the detector 3 and the domain of the training data of the prompt. The machine learning model that performs domain classification may be trained using any algorithm.

[0080] The domain determination unit 15B may determine the domain by receiving a domain class as an input to the system, or may determine the domain based on information input by the user via the input unit 40A.

[0081] The storage unit 20A also stores parameter sets corresponding to each of a plurality of domains. Each parameter set includes model parameters (at least a portion of the detector parameter θ) and / or additional model parameters updated by the detector parameter update unit 14A. The domain determination unit 15B selects a parameter set corresponding to the domain of the speech data from the plurality of model parameter sets. The parameter set selected by the domain determination unit 15B is used to detect artificial speech, as described below.

[0082] (Processing Unit) The processing unit 16B processes the speech data to be detected using the prompt p selected by the domain determination unit 15B, and supplies the processed speech data to the prediction unit 17B. As an example, the processing unit 16B divides the speech data into multiple parts, generates tokens representing features from each of the parts, and processes the generated tokens using the prompt p. On the other hand, when the source domain is determined by the domain determination unit 15B, the processing unit 16B supplies the tokens generated from the speech data to the prediction unit 17B without using a prompt.

[0083] (Prediction Unit) The prediction unit 17B inputs the voice data processed by the processing unit 16B to the detector 3, thereby determining whether the voice represented by the voice data to be detected is an artificial voice. The detector 3 used by the prediction unit 17B is a detector determined by the detector parameter θ that has not been updated by the detector parameter update unit 14A and a parameter set corresponding to the domain determined by the domain determination unit 15B. In other words, the prediction unit 17B performs a determination process by inputting the token processed by the processing unit 16B to the detector 3 determined by the model parameters and / or additional model parameters corresponding to the domain determined by the domain determination unit 15B.

[0084] (Flow of prediction method) Fig. 11 is a flow diagram showing an example of the flow of a prediction method S1B executed by the information processing device 1B. Some of the steps included in the flow diagram of Fig. 11 may be executed in parallel or in a different order. In S201, the domain determination unit 15B acquires voice data as input data. In addition, in step S202, the domain determination unit 15B acquires a trained detector 3. Note that acquiring the detector 3 means acquiring a detector parameter θ.

[0085] In step S203, the domain determination unit 15B determines the domain of the voice data. In step S204, the domain determination unit 15B selects a prompt p that corresponds to the domain of the voice data from among the plurality of prompts p.

[0086] In step S205, the domain determination unit 15B selects model parameters corresponding to the domain of the voice data from among the model parameters updated by the prompt update unit 13A. If the prompt update unit 13A has updated additional model parameters, in step S205 the domain determination unit 15B selects additional model parameters corresponding to the domain of the voice data from among the additional model parameters updated by the prompt update unit 13A.

[0087] In S206, the processing unit 16B processes the voice data using the prompt p determined in step S204. In S207, the prediction unit 17B determines whether the voice represented by the voice data is an artificial voice based on the result obtained by inputting the voice data processed in step S206 to the detector 3. In other words, the prediction unit 17B determines whether the input voice is an artificial voice by inputting the voice data processed by the processing unit 16B to the detector 3 determined by the model parameters selected by the domain determination unit 15B. Also, if the domain determination unit 15B selected additional model parameters in step S205, the prediction unit 17B can also determine whether the input voice is an artificial voice by inputting the voice data processed by the processing unit 16B to the detector 3 determined by the additional model parameters selected by the domain determination unit 15B.

[0088] (Effects of Information Processing Device) As described above, the information processing device 1B employs a configuration including a processing unit 16B that processes the voice data to be detected using a prompt p updated by the prompt updating unit 13A, and a prediction unit 17B that determines whether the voice represented by the voice data is an artificial voice based on a result obtained by inputting the voice data processed by the processing unit 16B into the detector 3. By using a prompt learned using training data of the target domain, the information processing device 1B can more accurately detect whether the voice represented by the voice data is an artificial voice.

[0089] The information processing device 1B further includes a domain determination unit 15B that selects a prompt p corresponding to the domain of the voice data from among the multiple prompts p updated by the prompt update unit 13A, and the processing unit 16B processes the voice data using the prompt p selected by the domain determination unit 15B. Therefore, the information processing device 1B can improve the detection accuracy for inputs from multiple domains.

[0090] Furthermore, in the information processing device 1B, a prediction unit 17B selects model parameters corresponding to the domain of the voice data from the model parameters updated by the detector parameter update unit 14A, and inputs the voice data processed by the processing unit 16B to a detector determined by the selected model parameters. Therefore, the information processing device 1B can improve the detection accuracy for inputs of multiple domains.

[0091] Furthermore, in the information processing device 1B, a prediction unit 17B selects additional model parameters corresponding to the domain of the voice data from the additional model parameters updated by the detector parameter update unit 14A, and inputs the voice data processed by the processing unit 16B to a detector determined by the selected additional model parameters. Therefore, the information processing device 1B can improve the detection accuracy for inputs of multiple domains.

[0092] Furthermore, the information processing device 1B employs a configuration in which the processing unit 16B generates tokens representing features from each of the multiple segments of divided voice data, and processes the generated tokens using the prompt updated by the prompt updating unit 13A. Therefore, according to the information processing device 1B, by inputting the tokens processed using the prompt to the detector 3, it is possible to more accurately detect whether the voice represented by the voice data is an artificial voice.

[0093] [Example of implementation by software] Some or all of the functions of the information processing devices 1, 1A, 1B (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.

[0094] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 12. Figure 12 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.

[0095] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.

[0096] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0097] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0098] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0099] [Additional Note 1] The functions of each of the above devices may be shared and implemented by multiple devices. For example, each of the above devices may be realized as a system in which multiple devices are connected via a communication network. In this case, as an example, the system may include a first device including a learning phase execution unit 111B and a second device including an estimation phase execution unit 112B.

[0100] Furthermore, the information processing device 1A or the information processing device 1B may be configured not to include the detector parameter update unit 14A. In this case, the model parameters and additional model parameters of the detector 3 are not updated, and only the prompt p is updated.

[0101] [Appendix 2] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims. [Appendix A] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.

[0102] (Appendix A1) An information processing device comprising: processing means for processing training data representing labeled speech using prompts; and update means for updating the prompts so as to reduce the difference between a predicted value obtained by inputting the training data processed by the processing means into a detector that detects artificial speech and the label assigned to the training data.

[0103] (Appendix A2) The information processing device according to Appendix A1, wherein the detector is a model generated by machine learning using training data belonging to a first domain, and the processing means processes training data belonging to a second domain different from the first domain using the prompt.

[0104] (Supplementary Note A3) The information processing device according to Supplementary Note A1 or A2, wherein the processing means executes at least one of a process of adding the prompt to the training data and a process of combining the prompt.

[0105] (Appendix A4) The information processing device according to any one of Appendices A1 to A3, wherein the processing means processes tokens representing respective feature amounts obtained by dividing the training data into a plurality of parts, using the prompt.

[0106] (Supplementary Note A5) The information processing device according to any one of Supplementary Notes A1 to A4, wherein the update means updates a model parameter of the detector in addition to the prompt so that a difference between the predicted value and the label becomes smaller.

[0107] (Supplementary Note A6) The information processing device according to any one of Supplementary Notes A1 to A5, wherein the update means updates additional model parameters of the detector in addition to the prompt so that a difference between the predicted value and the label becomes smaller.

[0108] (Appendix A7) The information processing device according to any one of Appendices A1 to A6, wherein the processing means selects a prompt corresponding to a domain of the training data from among a plurality of prompts, and processes the training data using the selected prompt.

[0109] (Supplementary Note A8) The information processing device according to Supplementary Note A2, wherein the processing means processes the training data belonging to the first domain in addition to the training data belonging to the second domain.

[0110] (Supplementary Note A9) The information processing device according to Supplementary Note A8, wherein the update means calculates a first difference between a first predicted value obtained from first training data belonging to the first domain and a label attached to the first training data, and a second difference between a second predicted value obtained from second training data belonging to the second domain and a label attached to the second training data, using a common loss function.

[0111] (Supplementary Note A10) The information processing device according to Supplementary Note A8, wherein the updating means updates the prompt using a weighted sum of a first difference between a first predicted value obtained from first training data belonging to the first domain and a label attached to the first training data, and a second difference between a second predicted value obtained from second training data belonging to the second domain and a label attached to the second training data.

[0112] (Appendix A11) The information processing device according to any one of Appendices A1 to A10, further comprising: a second processing means for processing the voice data to be detected using the prompt updated by the update means; and a determination means for determining whether the voice represented by the voice data is an artificial voice based on a result obtained by inputting the voice data processed by the second processing means into the detector.

[0113] (Appendix A12) The information processing device according to Appendix A11, further comprising a selection means for selecting a prompt corresponding to a domain of the voice data from among a plurality of prompts updated by the update means, wherein the second processing means processes the voice data using the prompt selected by the selection means.

[0114] (Appendix A13) The information processing device according to Appendix A12, wherein the selection means selects model parameters corresponding to a domain of the voice data from among the model parameters updated by the update means, and the discrimination means inputs the voice data processed by the second processing means to a detector determined by the model parameters selected by the selection means.

[0115] (Appendix A14) The information processing device according to appendix A12 or A13, wherein the selection means selects an additional model parameter corresponding to a domain of the voice data from the additional model parameters updated by the update means, and the discrimination means inputs the voice data processed by the second processing means to a detector determined by the additional model parameter selected by the selection means.

[0116] (Supplementary Note A15) The information processing device according to any one of Supplementary Notes A11 to A14, wherein the second processing means generates tokens representing features from each of the multiple parts into which the voice data is divided, and processes the generated tokens using the prompt updated by the update means.

[0117] (Appendix A16) An information processing device comprising: (i) a processing means for processing speech data to be detected using a prompt that has been updated so as to reduce the difference between a predicted value obtained by inputting training data representing labeled speech, the training data being processed using a prompt, into a detector that detects artificial speech, and (ii) the label assigned to the training data; and a determination means for determining whether the speech represented by the speech data to be detected is an artificial speech by inputting the speech data processed by the processing means into the detector.

[0118] [Appendix B] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0119] (Appendix B1) An information processing method including: a processing step in which at least one processor processes training data representing labeled speech using a prompt; and an update step in which the at least one processor updates the prompt so as to reduce a difference between a predicted value obtained by inputting the training data processed in the processing step into a detector that detects artificial speech and the label assigned to the training data.

[0120] (Appendix B2) The information processing method described in Appendix B1, wherein the detector is a model generated by machine learning using training data belonging to a first domain, and in the processing, the at least one processor processes training data belonging to a second domain different from the first domain using the prompt.

[0121] (Supplementary Note B3) The information processing method according to Supplementary Note B1 or B2, wherein in the processing, the at least one processor executes at least one of a process of adding the prompt to the training data and a process of combining the prompt.

[0122] (Appendix B4) The information processing method according to any one of Appendices B1 to B3, wherein in the processing, the at least one processor processes tokens representing each of a plurality of feature amounts obtained by dividing the training data using the prompt.

[0123] (Supplementary Note B5) The information processing method according to any one of Supplementary Notes B1 to B4, wherein in the updating process, the at least one processor updates model parameters of the detector in addition to the prompt so that a difference between the predicted value and the label becomes smaller.

[0124] (Supplementary Note B6) The information processing method according to any one of Supplementary Notes B1 to B5, wherein in the updating process, the at least one processor updates additional model parameters of the detector in addition to the prompt so that a difference between the predicted value and the label becomes smaller.

[0125] (Appendix B7) The information processing method according to any one of Appendices B1 to B6, wherein in the processing, the at least one processor selects a prompt corresponding to a domain of the training data from among a plurality of prompts, and processes the training data using the selected prompt.

[0126] (Supplementary Note B8) The information processing method according to Supplementary Note B2, wherein in the processing, the at least one processor processes training data belonging to the first domain in addition to training data belonging to the second domain.

[0127] (Supplementary Note B9) The information processing method according to Supplementary Note B8, wherein in the update process, the at least one processor calculates a first difference between a first predicted value obtained from first training data belonging to the first domain and a label attached to the first training data, and a second difference between a second predicted value obtained from second training data belonging to the second domain and a label attached to the second training data, using a common loss function.

[0128] (Supplementary Note B10) The information processing method described in Supplementary Note B8, wherein in the update process, the at least one processor updates the prompt using a weighted sum of a first difference between a first predicted value obtained from first training data belonging to the first domain and a label attached to the first training data, and a second difference between a second predicted value obtained from second training data belonging to the second domain and a label attached to the second training data.

[0129] (Appendix B11) The information processing method according to any one of Appendices B1 to B10, further comprising: a second processing step in which the at least one processor processes the voice data to be detected using the prompt updated in the update step; and a determination step in which the at least one processor determines whether the voice represented by the voice data is an artificial voice based on a result obtained by inputting the voice data processed in the second processing step into the detector.

[0130] (Appendix B12) The information processing method described in Appendix B11 further includes a selection process in which the at least one processor selects a prompt corresponding to a domain of the voice data from among a plurality of prompts updated in the update process, and the second processing process processes the voice data using the prompt selected in the selection process.

[0131] (Appendix B13) The information processing method described in Appendix B12, wherein in the selection process, the at least one processor selects model parameters corresponding to the domain of the voice data from the model parameters updated in the update process, and in the discrimination process, the at least one processor inputs the voice data processed in the second processing process to a detector determined by the model parameters selected in the selection process.

[0132] (Appendix B14) The information processing method described in Appendix B12 or B13, wherein in the selection process, the at least one processor selects additional model parameters corresponding to the domain of the audio data from the additional model parameters updated in the update process, and in the discrimination process, the at least one processor inputs the audio data processed in the second processing process to a detector determined by the additional model parameters selected in the selection process.

[0133] (Appendix B15) The information processing method according to any one of Appendices B11 to B14, wherein the second processing step generates tokens representing features from each of the multiple parts of the speech data that have been divided, and processes the generated tokens using the prompt updated in the update step.

[0134] (Appendix B16) An information processing method including: (i) a processing process in which at least one processor processes speech data to be detected using a prompt that has been updated so as to reduce the difference between a predicted value obtained by inputting training data representing labeled speech, the training data being processed using a prompt, into a detector that detects artificial speech, and (ii) the label assigned to the training data; and (ii) a discrimination process in which the at least one processor inputs the speech data processed in the processing process into the detector, and determines whether the speech represented by the speech data to be detected is an artificial speech.

[0135] [Appendix C] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0136] (Appendix C1) An information processing program that causes a computer to function as an information processing device, the information processing program causing the computer to function as: a processing means that processes training data representing labeled speech using prompts; and an update means that updates the prompts so as to reduce the difference between a predicted value obtained by inputting the training data processed by the processing means into a detector that detects artificial speech and the label assigned to the training data.

[0137] (Appendix C2) The information processing program according to Appendix C1, wherein the detector is a model generated by machine learning using training data belonging to a first domain, and the processing means processes training data belonging to a second domain different from the first domain using the prompt.

[0138] (Supplementary Note C3) The information processing program according to Supplementary Note C1 or C2, wherein the processing means executes at least one of a process of adding the prompt to the training data and a process of combining the prompt.

[0139] (Supplementary Note C4) The information processing program according to any one of Supplementary Notes C1 to C3, wherein the processing means processes tokens representing respective feature amounts obtained by dividing the training data into a plurality of parts, using the prompt.

[0140] (Supplementary Note C5) The information processing program according to any one of Supplementary Notes C1 to C4, wherein the updating means updates a model parameter of the detector in addition to the prompt so that a difference between the predicted value and the label becomes smaller.

[0141] (Supplementary Note C6) The information processing program according to any one of Supplementary Notes C1 to C5, wherein the updating means updates an additional model parameter of the detector in addition to the prompt so that a difference between the predicted value and the label becomes smaller.

[0142] (Appendix C7) The information processing program according to any one of Appendices C1 to C6, wherein the processing means selects a prompt corresponding to a domain of the training data from among a plurality of prompts, and processes the training data using the selected prompt.

[0143] (Supplementary Note C8) The information processing program according to Supplementary Note C2, wherein the processing means processes training data belonging to the first domain in addition to training data belonging to the second domain.

[0144] (Appendix C9) The information processing program according to Appendix C8, wherein the update means calculates a first difference between a first predicted value obtained from first training data belonging to the first domain and a label attached to the first training data, and a second difference between a second predicted value obtained from second training data belonging to the second domain and a label attached to the second training data, using a common loss function.

[0145] (Supplementary Note C10) The information processing program according to Supplementary Note C8, wherein the updating means updates the prompt using a weighted sum of a first difference between a first predicted value obtained from first training data belonging to the first domain and a label attached to the first training data, and a second difference between a second predicted value obtained from second training data belonging to the second domain and a label attached to the second training data.

[0146] (Appendix C11) An information processing program according to any one of appendices C1 to C10, further comprising: a second processing means for processing the voice data to be detected using the prompt updated by the update means; and a discrimination means for determining whether the voice represented by the voice data is an artificial voice based on the result obtained by inputting the voice data processed by the second processing means into the detector.

[0147] (Appendix C12) The information processing program described in Appendix C11, further causing the computer to function as a selection means for selecting a prompt corresponding to the domain of the voice data from among the plurality of prompts updated by the update means, and the second processing means processes the voice data using the prompt selected by the selection means.

[0148] (Appendix C13) The information processing program described in Appendix C12, wherein the selection means selects model parameters corresponding to the domain of the voice data from the model parameters updated by the update means, and the discrimination means inputs the voice data processed by the second processing means to a detector determined by the model parameters selected by the selection means.

[0149] (Appendix C14) The information processing program described in Appendix C12 or C13, wherein the selection means selects an additional model parameter corresponding to the domain of the voice data from the additional model parameters updated by the update means, and the discrimination means inputs the voice data processed by the second processing means to a detector determined by the additional model parameter selected by the selection means.

[0150] (Appendix C15) The information processing program according to any one of Appendices C11 to C14, wherein the second processing means generates tokens representing features from each of the multiple parts into which the voice data is divided, and processes the generated tokens using the prompt updated by the update means.

[0151] (Appendix C16) An information processing program that causes a computer to function as an information processing device, the information processing program causing the computer to function as: (i) a processing means that processes speech data to be detected using a prompt that has been updated so as to reduce the difference between a predicted value obtained by inputting training data representing labeled speech, the training data being processed using a prompt, into a detector that detects artificial speech, and (ii) the label assigned to the training data; and a discrimination process that determines whether the speech represented by the speech data to be detected is an artificial speech by inputting the speech data processed by the processing means into the detector.

[0152] 1, 1A, 1B Information processing device 3 Detector 4 Feature extractor 11A, 16B Processing unit 12A Loss calculation unit 13A Prompt update unit 14A Detector parameter update unit 15B Domain determination unit 17B Prediction unit

Claims

1. A processing means for processing training data representing labeled speech using prompts, An update means updates the prompt so that the difference between the predicted value obtained by inputting the training data processed by the processing means into a detector that detects artificial voice and the label attached to the training data becomes small. An information processing device equipped with the following features.

2. The detector is a model generated by machine learning using training data belonging to the first domain. The processing means processes training data belonging to a second domain different from the first domain using the prompt. The information processing apparatus according to claim 1.

3. The processing means performs at least one of the following processes: adding the prompts to the training data and combining them. The information processing apparatus according to claim 1 or 2.

4. The processing means processes tokens representing each of the features obtained by dividing the training data into multiple parts, using the prompt. The information processing apparatus according to claim 1 or 2.

5. The update means updates the model parameters of the detector in addition to the prompt so that the difference between the predicted value and the label becomes smaller. The information processing apparatus according to claim 1 or 2.

6. The update means updates the additional model parameters of the detector in addition to the prompt so that the difference between the predicted value and the label becomes smaller. The information processing apparatus according to claim 1 or 2.

7. The processing means selects a prompt from among a plurality of prompts that corresponds to the domain of the training data, and processes the training data using the selected prompt. The information processing apparatus according to claim 1 or 2.

8. The processing means processes the training data belonging to the first domain in addition to the training data belonging to the second domain. The information processing apparatus according to claim 2.

9. At least one processor processes training data representing labeled speech using prompts, The at least one processor performs an update process to update the prompt so that the difference between the predicted value obtained by inputting the processed training data into a detector for detecting artificial speech and the label attached to the training data becomes smaller. Information processing methods including

10. An information processing program that causes a computer to function as an information processing device, wherein the computer A processing means for processing training data representing labeled speech using prompts, An update means updates the prompt so that the difference between the predicted value obtained by inputting the training data processed by the processing means into a detector that detects artificial voice and the label attached to the training data becomes small. An information processing program designed to function as such.