Information processing system, information processing method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2026-03-25
AI Technical Summary
Existing machine translation techniques often generate translations that are excessively long for the environment in which they are output, leading to issues such as lyrics not fitting within the singing time of a song or speech bubbles in manga, making proper placement impossible.
An information processing method that generates instruction information including a first string and a length constraint, using a trained generative model to convert the string into a second language while adhering to the specified length constraints, such as singing time, speech duration, or character area size.
Reduces the likelihood of generating translations that are excessively long or short for the output environment, ensuring the translated text fits within the constraints of the original media, such as song duration, speech period, or speech bubble size.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to translation technology. [Background technology]
[0002] Machine translation techniques for converting a string of characters expressed in a specific language into another language have been proposed. For example, Patent Document 1 discloses a technique for generating translations of lyrics using translations based on song information such as the music genre or the gender of the singer. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2001-075963 Summary of the Invention [Problem to be solved by the invention]
[0004] However, simply using translations based on the music genre, the gender of the singer, or the like may result in, for example, an excessively long translation being generated for a song, which may be inappropriate as lyrics to be sung in parallel with the performance of the song. While the above description focuses on translating lyrics, similar problems are anticipated when translating character strings other than lyrics. For example, when translating lines placed in a speech bubble in a manga, a translation that is excessively long for the size of the speech bubble may be generated, resulting in an inability to properly position the translation in the speech bubble. In consideration of the above circumstances, one aspect of the present disclosure aims to reduce the possibility of generating a short string that is excessively long for the environment in which the converted character string is output. [Means for solving the problem]
[0005] In order to solve the above problems, an information processing method according to one embodiment of the present disclosure generates instruction information including a first string that expresses lyrics of a song, audio of content, or dialogue from a manga in a first language, and a constraint on the length of the string, and obtains a second string generated by a trained generative model processing the instruction information, the second string being obtained by converting the first string into a second language different from the first language under the constraint.
[0006] An information processing system according to one embodiment of the present disclosure includes an information generation unit that generates instruction information including a first string of characters that expresses lyrics of a song, audio of content, or dialogue from a manga in a first language and a constraint on the length of the string, and an information acquisition unit that acquires a second string that is generated by a trained generative model processing the instruction information and that is converted from the first string into a second language different from the first language under the constraint.
[0007] A program according to one embodiment of the present disclosure causes a computer system to function as an information generation unit that generates instruction information including a first string of characters that expresses lyrics of a song, audio of content, or dialogue from a manga in a first language and a constraint on the length of the string, and an information acquisition unit that acquires a second string that is generated by a trained generative model processing the instruction information and that is converted from the first string into a second language different from the first language under the constraint. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram of an information system according to a first embodiment. [Figure 2] FIG. 2 is a block diagram of a terminal device. [Figure 3] FIG. 2 is a block diagram illustrating a functional configuration of a terminal device. [Figure 4] 10 is a flowchart of a translation process. [Figure 5] FIG. 10 is a block diagram illustrating an example of the functional configuration of a terminal device according to a second embodiment. [Figure 6] 10 is a flowchart of a translation process in the second embodiment. [Figure 7] FIG. 11 is a block diagram illustrating a functional configuration of a terminal device according to a third embodiment. [Figure 8] FIG. 1 is an explanatory diagram illustrating the relationship between images and dialogue in manga data. [Figure 9] 10 is a flowchart of a translation process in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] A: First embodiment FIG. 1 is a block diagram illustrating the configuration of an information system 100 in the first embodiment. The information system 100 includes a terminal device 10 and a translation system 20. The terminal device 10 is an information processing system used by a user U. Examples of the terminal device 10 include portable information devices such as a smartphone, a tablet terminal, or a personal computer. The terminal device 10 can communicate with the translation system 20 via a communication network 200 such as the Internet.
[0010] 2 is a block diagram of the terminal device 10. The terminal device 10 includes a control device 11, a storage device 12, a communication device 13, an operation device 14, a display device 15, and a sound emitting device 16. The terminal device 10 may be realized as a single device, or may be realized as a plurality of devices configured separately from each other.
[0011] The control device 11 is composed of one or more processors that control each element of the terminal device 10. For example, the control device 11 is composed of one or more types of processors, such as a CPU (Central Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit). The communication device 13 communicates with the translation system 20 under the control of the control device 11.
[0012] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is configured with a known storage medium such as a magnetic storage medium or a semiconductor storage medium. The storage device 12 may be configured with a combination of multiple types of storage media. Furthermore, a portable storage medium that is detachable from the terminal device 10, or a storage medium (e.g., cloud storage) that the control device 11 can write to or read from via the communication network 200 may be used as the storage device 12.
[0013] The operation device 14 is an input device that receives instructions from the user U. The operation device 14 is, for example, an operator operated by the user U, or a touch panel that detects contact by the user U. The display device 15 displays images under the control of the control device 11. The display device 15 is configured with a display panel such as a liquid crystal panel or an organic EL panel. The sound emitting device 16 reproduces sound waves under the control of the control device 11. The sound emitting device 16 is, for example, a speaker or headphones.
[0014] 3 is a block diagram illustrating an example of the functional configuration of the terminal device 10. The control device 11 executes a program stored in the storage device 12 to realize multiple functions (a condition specifying unit 31, an information generating unit 32, an information acquiring unit 33, and an output control unit 34).
[0015] The storage device 12 of the first embodiment stores music data X1. The music data X1 is data representing the musical score of a music piece. The music data X1 is, for example, a music file that complies with the MIDI (Musical Instrument Digital Interface) standard. Specifically, the music data X1 includes a musical note sequence N and lyrics L1 of the music piece. The musical note sequence N is a time series of multiple notes that make up the music piece. The music data X1 specifies the pitch and sound duration for each note that makes up the musical note sequence N.
[0016] The lyrics L1 are a time sequence of characters corresponding to each note in the note sequence N. That is, the lyrics L1 correspond in time to the note sequence N. For example, one syllable constituting the lyrics L1 is assigned to one note in the note sequence N. The lyrics L1 in the first embodiment are a character string expressed in a first language. The first language is, for example, Japanese.
[0017] The translation system 20 in FIG. 1 is a machine translation device that converts lyrics L1 expressed in a first language into lyrics L2 expressed in a second language. The second language is different from the first language. The second language is, for example, English. That is, the translation system 20 generates lyrics L2 by converting the language of lyrics L1. Note that lyrics L1 is an example of a "first character string," and lyrics L2 is an example of a "second character string."
[0018] The condition specification unit 31 in FIG. 3 specifies a constraint Cx from the music data X1. The constraint Cx is a condition related to the length of the lyrics L2 after translation. In the first embodiment, the constraint Cx is set according to the duration of each syllable that constitutes the lyrics L1. Specifically, the condition specification unit 31 sets the constraint Cx according to the total duration of the multiple syllables that constitute the lyrics L1 (hereinafter referred to as "singing time"). The duration of each syllable is the duration of the pronunciation period of the note to which that syllable is assigned in the sequence of notes N. For example, the constraint Cx is set to the number of characters that can be uttered within the singing time of the lyrics L1. As explained above, the constraint Cx is a time condition related to the lyrics L2.
[0019] As explained above, the note sequence N in the music data X1 is accompanying information for specifying the constraint Cx. The accompanying information is a different type of information from the lyrics L1 and is provided in association with the lyrics L1. The note sequence N in the first embodiment can also be expressed as accompanying information used for playing the lyrics L1, accompanying information played together with the lyrics L1, or accompanying information that defines the conditions for playing the lyrics L1. On the other hand, the note sequence N (accompanying information) is not information that directly specifies the length of the lyrics L2 or information dedicated to setting the length of the lyrics L2.
[0020] The information generating unit 32 generates instruction information P. The instruction information P is a prompt that instructs the translation of lyrics L1. Specifically, the information generating unit 32 generates instruction information P that includes lyrics L1 and constraints Cx. For example, the instruction information P is a natural language prompt that requests that lyrics L1 be converted from a first language to a second language under constraints Cx. For example, the instruction information P is a prompt that expresses a request such as "Please translate Japanese lyrics (lyrics L1) into English under (constraints Cx)."
[0021] The information generation unit 32 transmits the instruction information P exemplified above from the communication device 13 to the translation system 20. The translation system 20 generates response information R by processing the instruction information P received from the terminal device 10, and transmits the response information R to the terminal device 10. The response information R is information including lyrics L2. The lyrics L2 are a character string obtained by converting lyrics L1 from the first language to the second language under constraints Cx. That is, for example, lyrics L2 in the second language are generated such that the singing time satisfies constraints Cx while retaining the meaning of lyrics L1.
[0022] The translation system 20 generates response information R by processing instruction information P using a trained generative model M. The generative model M is a generative probabilistic model that generates response information R in response to instruction information P. The generative model M has learned the tendency of response information R in response to instruction information P through prior machine learning (pre-training).
[0023] Specifically, the generative model M is an interactive large-scale language model (LLM) that is trained specifically for natural language processing tasks such as machine translation. For example, a natural language processing model realized by a transformer model using a self-attention mechanism is an example of the generative model M.
[0024] As shown in the above example, the response information R (lyrics L2) is a character string generated by the trained generative model M processing the instruction information P. Specifically, the lyrics L2 are a character string in a second language expressed with a singing time that satisfies the constraint Cx for the singing time of the lyrics L1. For example, lyrics L2 with the same singing time as the lyrics L1 are generated. Therefore, even if the lyrics L1 are the same, the longer the singing time of the lyrics L1, the more characters there are in the lyrics L2.
[0025] The information acquisition unit 33 acquires the response information R (lyrics L2). Specifically, the information acquisition unit 33 receives the response information R transmitted from the translation system 20 via the communication device 13. The information acquisition unit 33 also generates new music data X2 by combining the note sequence N of the existing music data X1 with the lyrics L2 of the response information R. As described above, the music data X2 is generated by replacing the lyrics L1 in the first language with the lyrics L2 in the second language while maintaining the note sequence N in the music data X1. The information acquisition unit 33 stores the music data X2 in the storage device 12.
[0026] The output control unit 34 causes the sound emitting device 16 to play the music data X1 or X2. For example, when a command to play the music data X1 is received from the user U, the output control unit 34 causes the sound emitting device 16 to play the music represented by the music data X1. The playback sound of the music data X1 is a mixture of the sounds played on an instrument corresponding to the note sequence N in the music data X1 and the singing sound of lyrics L1 in a first language sung at the pitch corresponding to the note sequence N. On the other hand, when a command to play the music data X2 is received from the user U, the output control unit 34 causes the sound emitting device 16 to play the music represented by the music data X2. The playback sound of the music data X2 is a mixture of the sounds played on an instrument corresponding to the note sequence N in the music data X2 and the singing sound of lyrics L2 in a second language sung at the pitch corresponding to the note sequence N. Any known voice synthesis technology may be used to generate the singing sound represented by the lyrics L1 or L2. The user U in the first embodiment may be a general consumer who listens to music, a producer who creates music (for example, a composer), or a seller who sells the music.
[0027] 4 is a flowchart of a process (hereinafter referred to as "translation process") executed by the control device 11. The translation process is started in response to an instruction from the user U via the operation device 14. When the translation process is started, the control device 11 (condition specification unit 31) specifies constraint conditions Cx from the music data X1 stored in the storage device 12 (Sa1).
[0028] The control device 11 (information generation unit 32) generates instruction information P including lyrics L1 and constraints Cx (Sa2) and transmits the instruction information P to the translation system 20 via the communication device 13 (Sa3). The translation system 20 processes the instruction information P using the generation model M to generate lyrics L2 and transmits response information R including the lyrics L2 to the terminal device 10.
[0029] The control device 11 (information acquisition unit 33) receives the response information R transmitted from the translation system 20 via the communication device 13 (Sa4). The control device 11 (information acquisition unit 33) generates music data X2 including the note sequence N of the music data X1 and the lyrics L2 of the response information R, and stores the music data X2 in the storage device 12 (Sa5).
[0030] As explained above, in the first embodiment, the generative model M processes the instruction information P, which includes the lyrics L1 to be converted as well as the constraint Cx, to generate lyrics L2, which are obtained by converting the lyrics L1 from the first language to the second language under the constraint Cx. This reduces the possibility of generating lyrics L2 that are excessively long or short for the sequence of notes N corresponding to the lyrics L1.
[0031] In the first embodiment, lyrics L2 are generated under a constraint Cx on singing time. This reduces the possibility of generating lyrics L2 with a singing time that deviates excessively from the singing time of lyrics L1. In other words, lyrics L2 in a second language that can be sung comfortably for an existing song can be generated. For example, lyrics L2 that can be sung within the playback time of the song can be generated.
[0032] In the first embodiment, the constraints Cx are identified from the music data X1, which includes the lyrics L1. That is, the music data X1 is used to identify both the lyrics L1 and the constraints Cx. Therefore, the configuration and processing required to generate the lyrics L2 can be simplified compared to an embodiment in which the constraints Cx are identified from elements separate from the music data X1.
[0033] B: Second embodiment 5 is a block diagram illustrating the functional configuration of the terminal device 10 in the second embodiment. The control device 11 executes a program stored in the storage device 12 to realize multiple functions (a condition specifying unit 31, an information generating unit 32, an information acquiring unit 33, and an output control unit 34).
[0034] The storage device 12 of the second embodiment stores content Y1. The content Y1 is a variety of media distributed or sold for viewing by a large number of viewers, and is composed of video V and audio A1. The video V is a live-action video captured by an imaging device, or a video such as an animation created using computer graphics. The audio A1 is sound played in parallel with the video V. In other words, the audio A1 corresponds in time to the video V. The audio A1 is spoken in a first language (e.g., Japanese). Note that the content Y1 may also be audio content (e.g., a podcast or audiobook) that does not include the video V.
[0035] The translation system 20 of the second embodiment converts a character string B1 in a first language expressed by a speech A1 into a character string B2 expressed in a second language. Note that the character string B1 is an example of a "first character string," and the character string B2 is an example of a "second character string."
[0036] The condition specification unit 31 specifies a constraint condition Cy from the content Y1. The constraint condition Cy is a condition related to the length of the translated character string B2. In the second embodiment, the constraint condition Cy is a condition related to the duration of the speech period when the character string B2 is spoken. The condition specification unit 31 sets the constraint condition Cy according to the duration of the speech period of the speech A1 (character string B1). For example, the constraint condition Cy is set to be the number of characters that can be spoken within the speech period of the speech A1.
[0037] The speech period of the speech A1 is a period during which the speech A1 exists in the content Y1. For example, a period during which speech is observed at a volume exceeding a predetermined threshold is exemplified as the speech period. Specifically, the speech period is the period from when playback of one phrase of the speech A1 starts to when playback of the phrase ends. Note that the period from when playback of one phrase of the speech A1 starts to when playback of the immediately following phrase starts may also be specified as the speech period. As explained above, the speech A1 of the content Y1 is associated information for specifying the constraint condition Cy.
[0038] The information generation unit 32 generates instruction information P. The instruction information P in the second embodiment is a prompt instructing the translation of a character string B1 represented by the speech A1. Specifically, the information generation unit 32 generates the instruction information P including the character string B1 and a constraint Cy. The instruction information P is a natural language prompt requesting that the character string B1 be converted from a first language to a second language under the constraint Cy. For example, the instruction information P is a prompt expressing a request such as "Please translate a Japanese sentence, (character string B1), into English under (constraint Cy)." The information generation unit 32 estimates the character string B1 by speech recognition of the speech A1 of the content Y1. Any known speech recognition technology may be employed to estimate the character string B1.
[0039] The information generation unit 32 transmits the instruction information P exemplified above from the communication device 13 to the translation system 20. The translation system 20 generates response information R by processing the instruction information P received from the terminal device 10, and transmits the response information R to the terminal device 10. The response information R is information including a character string B2. The character string B2 is a character string obtained by converting the character string B1 from the first language to the second language under the constraint condition Cy. That is, for example, a character string B2 in the second language is generated that retains the meaning of the character string B1 and whose speaking duration satisfies the constraint condition Cy.
[0040] As shown in the above example, the response information R (character string B2) is a character string generated by the trained generative model M processing the instruction information P. Specifically, the character string B2 is a character string in a second language that can be spoken with a duration that satisfies the constraint condition Cy with respect to the duration of the speech period in the content Y1. For example, the character string B2 is generated to be spoken with a duration equivalent to the speech period of the speech A1. Therefore, even if the character string B1 is the same, the longer the duration of the speech period of the speech A1, the more characters there are in the character string B2.
[0041] The information acquisition unit 33 acquires response information R (character string B2). Specifically, the information acquisition unit 33 receives the response information R transmitted from the translation system 20 via the communication device 13. The information acquisition unit 33 also generates new content Y2 by combining a video V of existing content Y1 with a voice A2 representing the character string B2 of the response information R. The voice A2 is a voice uttering the character string B2. Any known voice synthesis technology may be used to generate the voice A2. As described above, the content Y2 is generated by replacing the voice A1 in the first language with the voice A2 in the second language while maintaining the video V in the content Y1. The information acquisition unit 33 stores the content Y2 in the storage device 12.
[0042] The output control unit 34 plays content Y1 or content Y2. For example, when a command to play content Y1 is received from user U, the output control unit 34 causes the display device 15 to play video V of content Y1 and the sound emitting device 16 to play audio A1 in the first language. On the other hand, when a command to play content Y2 is received from user U, the output control unit 34 causes the display device 15 to play video V of content Y2 and the sound emitting device 16 to play audio A2 in the second language. The user U in the second embodiment may be a general consumer who views the content (Y1, Y2), as well as a producer who creates the content or a seller who sells the content.
[0043] 6 is a flowchart of the translation process in the second embodiment. When the translation process starts, the control device 11 (condition specifying unit 31) specifies constraint conditions Cy from the content Y1 stored in the storage device 12 (Sb1).
[0044] The control device 11 (information generation unit 32) estimates a character string B1 by speech recognition of the speech A1 (Sb2). The control device 11 (information generation unit 32) generates instruction information P including the character string B1 and a constraint condition Cy (Sb3) and transmits the instruction information P to the translation system 20 via the communication device 13 (Sb4). The translation system 20 processes the instruction information P using the generative model M to generate a character string B2 and transmits response information R including the character string B2 to the terminal device 10.
[0045] The control device 11 (information acquisition unit 33) receives the response information R transmitted from the translation system 20 via the communication device 13 (Sb5). The control device 11 (information acquisition unit 33) generates speech A2 by performing speech synthesis processing on the character string B2 (Sb6). The control device 11 (information acquisition unit 33) generates content Y2 including the video V of content Y1 and the synthesized speech A2, and stores the content Y2 in the storage device 12 (Sb7).
[0046] As described above, in the second embodiment, the instruction information P including the string B2 to be converted as well as the constraint Cy is processed by the generative model M, whereby the string B2 is generated by converting the string B1 from the first language to the second language under the constraint Cy. This reduces the possibility that an excessively long or short string B2 is generated for the video V corresponding to the string B1.
[0047] In particular, in the second embodiment, the character string B2 is generated under a constraint Cy regarding the duration of the speech period in the content Y1. Therefore, it is possible to reduce the possibility that an excessively long character string B2 is generated for the video V of the content Y1. In other words, it is easy to temporally correspond the playback of the video V and the playback of the audio A2 (character string B2) in the content Y1.
[0048] In the second embodiment, the constraint Cy is identified from the content Y1 including the audio A1. That is, the content Y1 is used to identify the character string B1 and the constraint Cy. Therefore, the configuration and processing required to generate the character string B2 can be simplified compared to a configuration in which the constraint Cy is identified from an element separate from the content Y1.
[0049] C-1: Third embodiment 7 is a block diagram illustrating the functional configuration of the terminal device 10 in the third embodiment. The control device 11 executes a program stored in the storage device 12 to realize a plurality of functions (a condition specifying unit 31, an information generating unit 32, an information acquiring unit 33, and an output control unit 34).
[0050] The storage device 12 of the third embodiment stores manga data Z1 representing a manga. The manga data Z1 is composed of an image G and dialogue W1. The dialogue W1 is a string of characters representing the words of a character appearing in the manga. The user U of the third embodiment may be a general consumer who reads manga, as well as a creator who creates the manga (e.g., a manga artist), or a publisher who publishes the manga.
[0051] FIG. 8 is an explanatory diagram of an image G and dialogue W1. A text area D is placed within the comic image G. The dialogue W1 is placed within the text area D. The text area D is a speech bubble placed in an appropriate size and shape within the comic image G. The dialogue W1 is a string of characters expressed in a first language (for example, Japanese).
[0052] The translation system 20 of the third embodiment converts a line W1 expressed in a first language into a line W2 expressed in a second language. Note that the line W1 is an example of a "first character string," and the line W2 is an example of a "second character string."
[0053] The condition specification unit 31 in FIG. 7 specifies a constraint Cz from the comic data Z1. The constraint Cz is a condition related to the length of the translated dialogue W2. In the third embodiment, the constraint Cz is a condition according to the size of the character area D in which the dialogue W1 is placed in the image G. The condition specification unit 31 specifies the size of the character area D in which the dialogue W1 is placed, and sets the constraint Cz according to the size of the character area D. For example, the constraint Cz for the dialogue W2 is set to be the number of characters that can be placed in the character area D of the dialogue W1. Any method can be used to specify the size of the character area D, but one example is to estimate the character area D in the image G by image recognition of the comic data Z1 and specify the size of the character area D. Furthermore, data indicating the size of the character area D may be included in the comic data Z1.
[0054] As explained above, image G (text area D) of comic data Z1 is accompanying information for identifying constraint condition Cz. The accompanying information is a different type of information from the dialogue W1 and is provided in association with the dialogue W1. Image G in the third embodiment can also be expressed as accompanying information used to display the dialogue W1, accompanying information displayed together with the dialogue W1, or accompanying information that specifies the conditions for displaying the dialogue W1. On the other hand, image G (text area D) is not information that directly specifies the length of the dialogue W2 or information dedicated to setting the length of the dialogue W2.
[0055] The information generating unit 32 generates instruction information P. The instruction information P in the third embodiment is a prompt that instructs the translation of the line W1 (i.e., conversion between languages). Specifically, the information generating unit 32 generates instruction information P that includes the line W1 and a constraint Cz. For example, the instruction information P is a natural language prompt that requests the line W1 to be converted from a first language to a second language under the constraint Cz. For example, the instruction information P is a prompt that expresses a request such as "Please translate the Japanese sentence (line W1) into English under (constraint Cz)."
[0056] The information generation unit 32 transmits the instruction information P exemplified above from the communication device 13 to the translation system 20. The translation system 20 generates response information R by processing the instruction information P received from the terminal device 10, and transmits the response information R to the terminal device 10. The response information R is information including a line W2. The line W2 is a character string obtained by converting the line W1 from the first language to the second language under the constraint Cz. That is, for example, a line W2 in the second language is generated that retains the meaning of the line W1 while having a number of characters that satisfies the constraint Cz.
[0057] As shown in the above example, the response information R (line W2) is a string generated by the trained generative model M processing the instruction information P. Specifically, the line W2 is a string in a second language expressed with a number of characters that satisfies the constraint Cz for the size of the character area D in which the line W1 is placed. For example, the line W2 is generated with a number of characters that can be placed in a character area D of the same size as the character area D of the line W1. Therefore, even if the line W1 is common, the larger the size of the character area D of the line W1, the more characters the line W2 has.
[0058] The information acquisition unit 33 acquires the response information R (the dialogue W2). Specifically, the information acquisition unit 33 receives the response information R transmitted from the translation system 20 via the communication device 13. The information acquisition unit 33 also generates new manga data Z2 by combining the image G of the existing manga data Z1 with the dialogue W2 of the response information R. As described above, the manga data Z2 is generated by replacing the dialogue W1 in the first language with the dialogue W2 in the second language while maintaining the image G in the manga data Z1. The information acquisition unit 33 stores the manga data Z2 in the storage device 12.
[0059] The output control unit 34 plays back the comic data Z1 or the comic data Z2. For example, when a command to play back the comic data Z1 is received from the user U, the output control unit 34 causes the display device 15 to display the image G of the comic data Z1 and the dialogue W1 in the first language arranged in the text area D. On the other hand, when a command to play back the comic data Z2 is received from the user U, the output control unit 34 causes the display device 15 to display the image G of the comic data Z2 and the dialogue W2 in the second language arranged in the text area D.
[0060] 9 is a flowchart of the translation process in the third embodiment. When the translation process starts, the control device 11 (condition specifying unit 31) specifies constraint conditions Cz from the comic data Z1 stored in the storage device 12 (Sc1).
[0061] The control device 11 (information generation unit 32) generates instruction information P including the lines W1 and the constraints Cz (Sc2) and transmits the instruction information P to the translation system 20 via the communication device 13 (Sc3). The translation system 20 processes the instruction information P using the generative model M to generate the lines W2 and transmits response information R including the lines W2 to the terminal device 10.
[0062] The control device 11 (information acquisition unit 33) receives the response information R sent from the translation system 20 via the communication device 13 (Sc4). The control device 11 (information acquisition unit 33) generates comic data Z2 including the image G of the comic data Z1 and the translated dialogue W2, and stores the comic data Z2 in the storage device 12 (Sc5).
[0063] As described above, in the third embodiment, the instruction information P including the line W2 to be converted as well as the constraint Cz is processed by the generative model M, whereby the line W2 is generated by converting the line W1 from the first language to the second language under the constraint Cz. This reduces the possibility that the line W2 is generated to be excessively long or short for the image G.
[0064] In particular, in the third embodiment, the line W2 is generated under a constraint Cz according to the size of the character area D in which the line W1 is placed. Therefore, it is possible to reduce the possibility of generating a line W2 that is excessively long for the size of the character area D. In other words, it is easy to place the line W2 in the character area D of the image G in the manga data Z1.
[0065] In the third embodiment, the constraints Cz are identified from the comic data Z1 that includes the lines W1. That is, the comic data Z1 is used to identify the lines W1 and the constraints Cz. Therefore, compared to a configuration in which the constraints Cz are identified from elements separate from the comic data Z1, the configuration and processing required to generate the lines W2 can be simplified.
[0066] As can be understood from the examples of the first to third embodiments, the music data X1 of the first embodiment, the content Y1 of the second embodiment, and the manga data Z1 of the third embodiment are collectively expressed as target data to be processed by the information system 100.
[0067] C-2: Modification of the third embodiment The third embodiment may be modified as shown below: According to the modifications as shown below, the same effects as those of the third embodiment can be achieved.
[0068] [Variation 1] There is a general tendency that the larger the number of characters in the line W1, the larger the size of the character area D in which the line W1 is placed. Taking the above tendency into consideration, the condition specification unit 31 may estimate the size of the character area D according to the number of characters in the line W1, and specify the constraint condition Cz according to the estimation result (i.e., the size of the character area D). For example, the condition specification unit 31 estimates the size of the character area D so that the larger the number of characters in the line W1, the larger the estimated size. The relationship between the number of characters in the line W1 and the size of the character area D is set statistically in advance, for example. The operation of setting the constraint condition Cz according to the size of the character area D is the same as in the third embodiment.
[0069] In the first modification, the size of the character area D is estimated according to the number of characters in the line W1. Therefore, the size of the character area D can be estimated by the simple process of counting the number of characters in the line W1.
[0070] [Variation 2] There is a general tendency that the larger the size of each character (hereinafter referred to as "character size") constituting the line W1, the larger the size of the character area D in which the line W1 is arranged. Taking the above tendency into consideration, the condition specification unit 31 may estimate the size of the character area D according to the number of characters and character size of the line W1, and specify the constraint condition Cz according to the estimation result. For example, the condition specification unit 31 estimates the size of the character area D so that the larger the number of characters in the line W1, the larger the estimated size of the character area D, and the larger the character size of the line W1, the larger the estimated size of the character area D. The relationship between the character size of the line W1 and the size of the character area D is set in advance, for example, statistically. The operation of setting the constraint condition Cz according to the size of the character area D is the same as in the third embodiment.
[0071] In variant example 2, the character size as well as the number of characters in the dialogue W1 is reflected in estimating the size of the character area D. Therefore, the size of the character area D can be estimated with high accuracy compared to a form in which the size of the character area D is estimated taking into account only the number of characters in the dialogue W1.
[0072] [Variation 3] The line W1 is composed of a plurality of characters arranged horizontally and vertically. For example, in a language written vertically (e.g., Japanese or Chinese), a set of a plurality of characters arranged vertically (one line) is arranged horizontally. In a language written horizontally (e.g., English), a set of a plurality of characters arranged horizontally (one line) is arranged vertically.
[0073] The size of the character area D correlates with the number of characters N1 in the vertical direction and the number of characters N2 in the horizontal direction of the line W1. For example, there is a tendency that the larger the number of characters N1 in the vertical direction, the larger the size of the character area D in the vertical direction, and the larger the number of characters N2 in the horizontal direction, the larger the size of the character area D in the horizontal direction. Taking these trends into consideration, the condition specification unit 31 estimates the size of the character area D according to the number of characters N1 in the vertical direction and the number of characters N2 in the horizontal direction of the line W1, and specifies the constraint Cz according to the estimation result (the size of the character area D). For example, the condition specification unit 31 estimates the size of the character area D so that the larger the number of characters N1 or N2 in the line W1, the larger the estimated size of the character area D. The relationship between the number of characters N1 and N2 in the line W1 and the size of the character area D is set statistically in advance, for example. The operation of setting the constraint Cz according to the size of the character area D is the same as in the third embodiment.
[0074] In variant example 3, the number of characters N1 in the vertical direction and the number of characters N2 in the horizontal direction of the dialogue W1 are reflected in estimating the size of the character area D. Therefore, even if there is a large difference between the number of characters N1 in the vertical direction and the number of characters N2 in the horizontal direction, the size of the character area D can be appropriately estimated.
[0075] [Variation 4] As mentioned above, the writing direction (vertical / horizontal) depends on the language. Specifically, there are languages in which vertical writing is applied (for example, Japanese or Chinese) and languages in which horizontal writing is applied (for example, English). In a text area D in which a line W1 is written vertically, horizontally written line W2 tends to be difficult to place. Similarly, in a text area D in which a line W1 is written horizontally, vertically written line W2 tends to be difficult to place. On the other hand, in a text area D in which a line W1 is written vertically, vertically written line W2 tends to be easily placed. Similarly, in a text area D in which a line W1 is written horizontally, horizontally written line W2 tends to be easily placed.
[0076] Taking the above tendency into consideration, the condition specification unit 31 of Modification 4 specifies the constraint condition Cz according to the relationship between the writing direction of the line W1 and the writing direction of the second character string. For example, when the writing direction is the same between the line W1 and the line W2, the number of characters specified by the constraint condition Cz exceeds the number of characters specified by the constraint condition Cz when the writing direction is different between the line W1 and the line W2. In other words, when the writing direction is the same between the line W1 and the line W2, the condition specification unit 31 sets the length of the line W2 specified by the constraint condition Cz to a large value. On the other hand, when the writing direction is different between the line W1 and the line W2, the condition specification unit 31 sets the length of the line W2 specified by the constraint condition Cz to a small value.
[0077] According to the fourth modification, the constraint Cz is set according to the relationship (for example, whether different or similar) between the writing direction (vertical / horizontal) of the line W1 and the writing direction of the line W2. Therefore, the line W2 can be generated with the number of characters that can be appropriately arranged in the character area D.
[0078] [Variation 5] For example, in a language such as English that is basically written in units of words, it would appear unnatural if a single word constituting the dialogue W2 were written across multiple lines in the character area D. In other words, it is preferable that each word of the dialogue W2 be arranged within a single line.
[0079] Taking the above circumstances into consideration, the constraint condition Cz of Modification 5 includes a condition (hereinafter referred to as an "auxiliary condition") regarding the number of characters in one word (hereinafter referred to as "word length") in the line W2. That is, the condition specification unit 31 of Modification 5 sets the constraint condition Cz including the auxiliary condition. The auxiliary condition specifies, for example, an upper limit value for word length. As explained above, the constraint condition Xz of Modification 5 includes a condition regarding the word length of the line W2 after translation, in addition to a condition regarding the length of the line W2.
[0080] Specifically, the condition specification unit 31 sets the auxiliary condition according to the size of the character area D. Specifically, the auxiliary condition is generated according to the size of the character area D so that the larger the size of the character area D, the larger the word length specified by the auxiliary condition. The relationship between the size of the character area D and the word length specified by the auxiliary condition is set statistically in advance, for example. The operation of the information generation unit 32 to generate instruction information P including the line W1 and the constraint condition Cz is the same as in the third embodiment. The instruction information P represents, for example, a request such as "Please translate the Japanese sentence (line W1) into English under (constraint condition Cz)."
[0081] As in the third embodiment, the translation system 20 processes the instruction information P to generate response information R. That is, a line W2 in a second language is generated that maintains the meaning of the line W1 while satisfying the constraint Cz in the number of characters. The line W2 is composed of words whose number of characters is equal to or less than the word length specified by the auxiliary information. Therefore, each word that composes the line W2 is displayed on a single line in the character area D, and a single word is not split across multiple lines.
[0082] As described above, according to the fifth modification, the constraint Cz includes an auxiliary condition related to the number of characters (word length) of one word in the line W2, making it difficult to generate a line W2 whose word length is extremely long relative to the character area D. In other words, it is possible to generate a line W2 that can be appropriately placed in the character area D.
[0083] [Variation 6] User U is the creator (e.g., a cartoonist) who creates the manga represented by manga data Z1, or the publisher who publishes the manga. By operating operation device 14 of terminal device 10, user U can instruct changes to the position and size of text area D in the created manga. For example, in manga image G, text area D (speech bubble) and the other drawing areas are generated as different layers. The drawing area is the area where the original illustrations of the manga are drawn. As described above, because text area D and the drawing area are set on separate layers, user U can instruct changes to only text area D while maintaining the drawing area.
[0084] The condition specification unit 31 changes the character area D of the image G represented by the comic data Z1 in response to an instruction from the user U. For example, the condition specification unit 31 changes the position and size of the character area D in response to an instruction from the user U. The condition specification unit 31 specifies the constraint condition Cz in response to the size of the character area D after the change. The operation of specifying the constraint condition Cz in response to the size of the character area D is the same as in the third embodiment.
[0085] As described above, according to the sixth modification, the user U can change the character area D of the comic in a variety of ways. Furthermore, since the constraint condition Cz is specified according to the size of the changed character area D, it is possible to generate lines W2 that can be appropriately placed in the changed character area D.
[0086] D: Modification Specific modified embodiments that can be added to each of the embodiments exemplified above are exemplified below. Two or more embodiments arbitrarily selected from the following examples may be combined as appropriate within the scope of not being mutually contradictory.
[0087] (1) The content of the constraints C (Cx, Cy, Cx) for translation using the generative model M is not limited to the examples in the above-mentioned embodiments. In the following description, a character string before translation expressed in a first language is referred to as a "first character string," and a character string in a second language generated by translating the first character string is referred to as a "second character string." The lyrics L1 of the first embodiment, the character string B1 (audio A1) of the second embodiment, and the line W1 of the third embodiment are examples of the "first character string," and the lyrics L2 of the first embodiment, the character string B2 (audio A2) of the second embodiment, and the line W2 of the third embodiment are examples of the "second character string."
[0088] [Aspect 1] Constraint condition C in aspect 1 includes, for example, a condition related to the size of the display surface of display device 15 (hereinafter referred to as "display size"). That is, condition specification unit 31 sets constraint condition C according to the display size of display device 15. For example, constraint condition C is exemplified as a number of characters that is suitable for the display size of display device 15. As a result of translating the first character string under constraint condition C exemplified above, a second character string is generated with a number of characters according to the display size. For example, even if the first character string is the same, the larger the display size, the greater the number of characters in the second character string.
[0089] [Aspect 2] Constraint condition C in aspect 2 includes a condition related to characters (hereinafter referred to as "characters") appearing in content such as video or manga. That is, condition specification unit 31 sets constraint condition C according to the characteristics of characters appearing in the content. The characteristics of characters appearing are, for example, features related to the character's personality or appearance. For example, they are specified by user U through an operation on operation device 14. For example, constraint condition C is exemplified as a number of characters that is appropriate for the dialogue of a character appearing. As a result of translating a first character string under constraint condition C exemplified above, a second character string is generated with a number of characters according to the characteristics of the character appearing. For example, even if the first character string is the same, if the characteristics of the character appearing differ, the number of characters in the second character string will also differ.
[0090] [Aspect 3] Constraint condition C in aspect 3 includes a condition related to the overall story of the content, such as a video or manga. That is, condition specification unit 31 sets constraint condition C according to the story of the content. The story of the content is specified by user U, for example, by operating operation device 14. For example, constraint condition C may be exemplified as a number of characters that is appropriate for the story of the content. As a result of translating the first character string under constraint condition C exemplified above, a second character string having a number of characters according to the story of the content is generated.
[0091] (2) In the above-described embodiments, the condition specification unit 31 specifies the constraints C (Cx, Cy, Cz) using the target data (music data X1, content Y1, manga data Z1). However, the method of specifying the constraints C is not limited to the above examples. For example, the condition specification unit 31 may set the constraints C in response to an instruction from the user U via the operation device 14. Specifically, the condition specification unit 31 specifies, as the constraints C, a condition selected by the user U from a plurality of conditions prepared in advance. The condition specification unit 31 may also specify, as the constraints C, a character string input by the user U via an operation on the operation device 14.
[0092] (3) In the second embodiment, content Y1 including video V and audio A1 is exemplified, but content Y1 may also be composed of video V and character string B1. Character string B1 is, for example, a subtitle displayed together with video V. In a form in which content Y1 includes character string B1, speech recognition (Sb2) for audio A1 is omitted.
[0093] Similarly, content Y2 may be composed of video V and character string B2. Character string B2 is, for example, a subtitle displayed together with video V. When content Y2 includes character string B2, speech synthesis (Sb6) for character string B2 is omitted.
[0094] (4) The target data to be processed in the third embodiment is not limited to the comic data Z1. The third embodiment is also applicable to various publications other than comics, such as magazines or newspapers. Therefore, the character area D in which the first character string or the second character string is arranged is not limited to a comic speech bubble. For example, a background area without a clear boundary may be identified as the character area D in which the first character string is arranged. Furthermore, an area in the comic image G in which an arbitrary object (e.g., a building or a natural object) is drawn may be identified as the character area D in which the first character string is arranged. Examples of configurations for identifying the character area D exemplified above include a configuration in which data representing the character area D is included in the comic data Z1, a configuration in which the character area D is identified by image recognition of the image G of the comic data Z1, or a configuration in which the user U indicates the character area D by operating the operation device 14.
[0095] (5) In the second embodiment, the condition specifying unit 31 specifies the constraint condition Cy from the audio A1 of the content Y1. However, the condition specifying unit 31 may specify the constraint condition Cy by analyzing the video V of the content Y1. For example, the condition specifying unit 31 sets the constraint condition Cy according to the length of the playback time of the video V. For example, the constraint condition Cy is set to be the number of characters that can be uttered within the playback time of the video V.
[0096] In the above configuration, the video V of the content Y1 is accompanying information for identifying the constraints Cy. The accompanying information is a different type of information from the audio A1 (character string B1) and is provided in association with the audio A1. The video V in the second embodiment can also be expressed as accompanying information used to play the audio A1, accompanying information played together with the audio A1, or accompanying information that specifies the conditions for playing the audio A1. On the other hand, the video V (accompanying information) is not information that directly specifies the length of the character string B2 or information dedicated to setting the length of the character string B2.
[0097] As can be understood from the above examples, the condition specification unit 31 in some embodiments of the present disclosure specifies the constraint condition C by using accompanying information included in the target data. The accompanying information is a comprehensive concept that includes the sequence of notes N in the music data X1 in the first embodiment, the video V or audio A1 in the content Y1 in the second embodiment, and the image G in the comic data Z1 in the third embodiment.
[0098] (6) In the second embodiment, the translation process may be executed in real time in parallel with the reproduction of the content Y1. For example, the control device 11 executes the translation process of Fig. 6 for each pronunciation period in parallel with the reproduction of the content Y1.
[0099] Furthermore, in the second embodiment, the constraint condition Cy may include a condition regarding information that is not to be included in the character string B2. For example, the condition specification unit 31 sets the constraint condition Cy including a condition that information detected from the video V of the content Y1 is not to be included in the character string B2. The information detected from the video V is, for example, information such as the name of an object included in the video V. Any known object detection technology may be employed to detect an object included in the video V. With the above configuration, the number of characters in the character string B2 is reduced, and as a result, the time required to play the audio A2 corresponding to the character string B2 is shortened.
[0100] (7) In the above-described embodiments, a first character string in a first language is converted into a second character string in a second language. However, the first language and the second language may be a common language. That is, the generative model M may convert a first character string expressed in a specific language into a second character string in that language under constraint C. When the number of characters in the second character string exceeds the number of characters in the first string, the processing by the generative model M corresponds to "supplementing" the first character string. When the number of characters in the second character string is less than the number of characters in the first string, the processing by the generative model M corresponds to "summarizing" the first character string. The above-described embodiments are similarly applicable to the supplementation or summarization of the first character string.
[0101] (8) In the above-described embodiments, the translation system 20 separate from the terminal device 10 generates the response information R. However, the terminal device 10 may generate the response information R. For example, the terminal device 10 may be equipped with a generative model M. The information acquisition unit 33 generates the response information R by processing the instruction information P generated by the information generation unit 32 using the generative model M. As can be understood from the above explanation, the information acquisition unit 33 is comprehensively expressed as an element that acquires the response information R. The "acquisition" of the response information R includes not only the "reception" of the response information R but also the "generation" of the response information R.
[0102] Furthermore, as can be understood from the above explanation, the “information processing system” in this disclosure may be the terminal device 10 alone in each of the above-mentioned forms, or the entire information system 100 including the terminal device 10 and the translation system 20.
[0103] (9) The generative model M is not limited to the large-scale language model exemplified in each of the above embodiments. For example, various models such as a recurrent neural network (RNN), a long short-term memory (LSTM), or a bidirectional encoder representations from transformers (BERT) can be used as the generative model M.
[0104] (10) The functions of the terminal device 10 according to each of the above-described embodiments are realized by cooperation between one or more processors constituting the control device 11 and a program stored in the storage device 12. The program according to each of the above-described embodiments can be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium is, for example, a non-transitory recording medium, such as an optical recording medium (optical disk) such as a CD-ROM, but also includes any known type of recording medium, such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium other than a transitory, propagating signal, and does not exclude volatile recording media. Furthermore, in a configuration in which a distribution device distributes a program via a communication network 200, the recording medium storing the program in the distribution device corresponds to the non-transitory recording medium described above.
[0105] E: Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.
[0106] An information processing method according to one aspect (aspect 1) of the present disclosure generates instruction information including a first string of characters that expresses lyrics of a song, audio of a content, or dialogue from a manga in a first language, and a constraint on the length of the string; and obtains a second string generated by a trained generative model processing the instruction information, the second string being obtained by converting the first string into a second language different from the first language under the constraint. According to the above aspect, by processing instruction information including a constraint on the length of the string in addition to the first string to be converted using a generative model, a second string is generated by converting the first language into the second language under the constraint. This reduces the possibility of generating a second string that is excessively long for the environment in which the second string is output.
[0107] A "trained generative model" is a probabilistic model that generates a second character string in response to instruction information. The generative model has learned the tendency of the second character string in response to instruction information including a first character string and constraints through prior machine learning. Specifically, the "generative model" is a generative large-scale language model (LMM) that has been trained specifically for natural language processing tasks such as response generation. For example, a probabilistic model realized by a Transformer model that utilizes a self-attention mechanism is an example of a "generative model."
[0108] A "constraint" is any condition related to the length of a character string. For example, a "constraint" may be a condition related to the total number of characters (number of characters) constituting the character string, a temporal condition such as the length of time the character string is pronounced, or a spatial condition such as the size at which the character string is displayed. In other words, the "length" of a character string is a concept that includes not only the number of characters, but also the length of time required for pronunciation or the size required for display.
[0109] In a specific example (Aspect 2) of Aspect 1, the first character string is the lyrics, and the constraint is a time constraint related to the second character string. According to the above aspect, lyrics are converted from a first language to a second language under a time constraint. Therefore, it is possible to reduce the possibility that a second character string that is excessively long or short for a sequence of notes in a piece of music is generated.
[0110] In a specific example (Aspect 3) of Aspect 2, the constraint is a condition according to the duration of each syllable in the first character string. According to the above aspect, the second character string is generated under a constraint according to the duration of each syllable. Therefore, it is possible to reduce the possibility of generating a second character string whose duration deviates excessively from the duration of each syllable in the first character string. In other words, it is possible to generate lyrics (second character string) in a second language that can be sung comfortably for an existing song.
[0111] In a specific example (Aspect 4) of Aspect 1, the first character string is a character string represented by the audio, and the constraint is a condition according to the duration of a speech period in the content. In the above aspect, the second character string is generated under a constraint according to the duration of a speech period in the content. This reduces the possibility of generating a second character string that is excessively long for the content. In other words, it is easy to temporally match the playback of video in the content with the playback of audio corresponding to the second character string.
[0112] In a specific example (Aspect 5) of Aspect 1, the first character string is a line from the comic, and the constraint is a condition according to the size of a character area in the comic where the first character string is placed. In the above aspect, the second character string is generated under a constraint according to the size of the character area where the line is placed. Therefore, it is possible to reduce the possibility of generating a second character string that is excessively long relative to the size of the character area. In other words, it is easy to place the second character string in the character area of the comic.
[0113] In a specific example (aspect 6) of aspect 5, the constraints are further identified from comic data including the first character string, and in generating the instruction information, the instruction information is generated including the first character string of the comic data and the constraints identified from the comic data. In the above aspect, the constraints are identified from comic data including the first character string. In other words, comic data is used both to identify the first character string and to identify the constraints. Therefore, the configuration and processing required to generate the second character string can be simplified compared to an embodiment in which constraints are identified from elements separate from the comic data.
[0114] In a specific example (aspect 7) of aspect 6, the constraint condition is identified by estimating the size of the character area based on the number of characters in the first character string, and identifying the constraint condition based on the size of the character area. In the above aspect, the size of the character area is estimated based on the number of characters in the first character string. Therefore, the size of the character area can be estimated by a simple process of counting the number of characters in the first character string.
[0115] In a specific example of aspect 7 (aspect 8), the size of the character region is estimated based on the number of characters and character size of the first character string. In the above aspect, the character size as well as the number of characters of the first character string is reflected in the estimation of the size of the character region. Therefore, the size of the character region can be estimated with higher accuracy than in an aspect in which the size of the character region is estimated based only on the number of characters of the first character string.
[0116] In a specific example (aspect 9) of any of aspects 6 to 8, when specifying the constraints, the size of the character area is estimated based on the number of characters in the vertical direction and the number of characters in the horizontal direction of the first character string, and the constraints are specified based on the size of the character area. According to the above aspects, the number of characters in the vertical direction and the number of characters in the horizontal direction of the first character string are reflected in the estimation of the size of the character area. Therefore, even if there is a large difference between the number of characters in the vertical direction and the number of characters in the horizontal direction, the size of the character area can be appropriately estimated.
[0117] In a specific example (aspect 10) of any of aspects 6 to 9, the constraint condition is specified according to the relationship between the writing direction of the first character string and the writing direction of the second character string. According to the above aspects, the constraint condition is set according to the relationship (for example, whether it is different or the same) between the writing direction of the first character string (vertical writing / horizontal writing) and the writing direction of the second character string, so that a second character string with a number of characters that can be appropriately arranged in a character area can be generated.
[0118] In a specific example (Aspect 11) of any of Aspects 6 to 10, when specifying the constraint conditions, an auxiliary condition regarding the number of characters in one word in the second character string is specified according to the size of the character area, and the constraint conditions including the auxiliary condition are specified. According to the above aspect, since the constraint conditions include an auxiliary condition regarding the number of characters in one word in the second character string, it becomes difficult to generate a second character string in which the number of characters in one word is excessively large compared to the character area. In other words, it is possible to generate a second character string with a number of characters that can be appropriately arranged in the character area.
[0119] In a specific example (aspect 12) of any of aspects 6 to 11, when specifying the constraint conditions, the character area is changed in accordance with instructions from the user, and the constraint conditions are specified according to the size of the changed character area. According to the above aspects, the user can change the character area in a comic in a variety of ways. Furthermore, because the constraint conditions are specified according to the size of the changed character area, a second string of characters can be generated that can be appropriately placed in the changed character area.
[0120] In a specific example (Aspect 13) of any of Aspects 1 to 12, the constraint condition is further identified from target data including the first character string, and the instruction information includes the first character string of the target data and the constraint condition identified from the target data. In the above aspects, the constraint condition is identified from target data including the first character string. That is, the target data is shared for identifying the first character string and the constraint condition. Therefore, compared to an aspect in which the constraint condition is identified from an element separate from the target data, the configuration and processing required to generate the second character string can be simplified.
[0121] An information processing system according to one aspect (aspect 14) of the present disclosure includes an information generation unit that generates instruction information including a first string of characters that expresses lyrics of a song, audio of content, or dialogue from a manga in a first language and a constraint on the length of the string, and an information acquisition unit that acquires a second string that is generated by a trained generative model processing the instruction information, and that is obtained by converting the first string into a second language different from the first language under the constraint.
[0122] A program according to one aspect (aspect 15) of the present disclosure causes a computer system to function as an information generation unit that generates instruction information including a first string of characters that expresses lyrics of a song, audio of content, or dialogue from a manga in a first language and a constraint on the length of the string, and an information acquisition unit that acquires a second string that is generated by a trained generative model processing the instruction information and that is converted from the first string into a second language different from the first language under the constraint. [Explanation of symbols]
[0123] 100...information system, 10...terminal device, 11...control device, 12...storage device, 13...communication device, 14...operation device, 15...display device, 16...sound emission device, 20...translation system, 31...condition specification unit, 32...information generation unit, 33...information acquisition unit, 34...output control unit.
Claims
1. An information generating unit that generates instruction information including a first character string that expresses a voice or character string of content or a line of a manga in a first language and a constraint on the length of the character string; an information acquisition unit that acquires a second character string generated by processing the instruction information, the second character string being obtained by converting the first character string into a second language different from the first language under the constraint condition; An information processing system comprising:
2. The second character string is a character string generated by a trained generative model processing the instruction information. The information processing system of claim 1.
3. The instruction information is a natural language prompt for the generative model.
3. The information processing system of claim 2.
4. The constraint is a condition regarding the number of characters in the character string.
4. The information processing system according to claim 1.
5. the first character string is a character string represented by the voice, The constraint is a condition according to the duration of a speech period in the content.
4. The information processing system according to claim 1.
6. the first character string is a line from the comic; The constraint is a condition according to the size of a character area in which the first character string is arranged in the comic.
4. The information processing system according to claim 1.
7. A condition specifying unit that specifies the constraint condition from manga data including the first character string, The information generating unit generates the instruction information including the first character string of the comic data and the constraint specified from the comic data.
7. The information processing system of claim 6.
8. The condition specifying unit A size of the character region is estimated according to the number of characters in the first character string, and the constraint condition is identified according to the size of the character region. The information processing system of claim 7.
9. The condition specification unit: The size of the character region is estimated according to the number of characters and the character size of the first character string. The information processing system of claim 8.
10. The condition specifying unit The size of the character area is estimated according to the number of characters in the vertical direction and the number of characters in the horizontal direction of the first character string, and the constraint condition is specified according to the size of the character area. The information processing system of claim 7.
11. The condition specifying unit The constraint is identified according to a relationship between a writing direction of the first character string and a writing direction of the second character string. The information processing system of claim 7.
12. The constraint is a condition regarding the number of characters in the character string, The number of characters specified by the constraint condition when the writing direction is the same between the first character string and the second character string is greater than the number of characters specified by the constraint condition when the writing direction is different between the first character string and the second character string. The information processing system of claim 11.
13. The condition specifying unit specifying an auxiliary condition regarding the number of characters in one word in the second character string according to the size of the character region; Identifying the constraints, including the auxiliary conditions The information processing system of claim 7.
14. The auxiliary condition specifies an upper limit on the number of characters in the word, the second character string is composed of one or more words having a number of characters equal to or less than the upper limit, Each of the one or more words is arranged in one line within the character area.
14. The information processing system of claim 13.
15. The condition specifying unit changing the character area in response to an instruction from a user; The constraint condition is identified according to the changed size of the character area. The information processing system of claim 7.
16. The character area is a speech bubble of the comic, a background area, or an area in which any object is drawn. The information processing system of claim 7.
17. The character area is identified based on data included in the manga data. The information processing system of claim 7.
18. The character area is identified by image recognition of the image of the manga data. The information processing system of claim 7.
19. The character area is specified by an instruction from a user to an operation device. The information processing system of claim 7.
20. A condition identification unit that identifies the constraint condition from target data including the first character string, The instruction information includes the first character string of the target data and the constraint specified from the target data.
4. The information processing system according to claim 1.
21. The constraint conditions include a condition according to the display size of the content or the manga.
4. The information processing system according to claim 1.
22. The constraint conditions include conditions according to the content or the characters or story of the manga.
4. The information processing system according to claim 1.
23. 1. A computer-implemented information processing method, comprising: an information generating step of generating instruction information including a first character string that expresses a voice or character string of the content or a line of a comic in a first language and a constraint on the length of the character string; an information acquiring step of acquiring a second character string generated by processing the instruction information, the second character string being obtained by converting the first character string into a second language different from the first language under the constraint condition; An information processing method including:
24. An information generation process for generating instruction information including a first character string that expresses the audio or character string of the content, or the dialogue of the manga in a first language, and a constraint on the length of the character string; an information acquiring step of acquiring a second character string generated by processing the instruction information, the second character string being obtained by converting the first character string into a second language different from the first language under the constraint condition; A program that causes a computer to execute the following.