Data processing apparatus, data processing method, and data processing program
The data processing apparatus enhances user immersion in electronic content by using a generation model to generate artificial voices with intonation and emotion based on character emotions, addressing the limitations of existing models.
Patent Information
- Application Number
- JP2024080476
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-05-16
AI Technical Summary
Existing generation models, such as language models, lack the capability to output high-quality artificial voices with intonation and emotion, leading to reduced user immersion in electronic content.
A data processing apparatus that inputs a prompt including an image of a specific section from electronic content and an instruction to estimate the emotion of a character, using a generation model to generate artificial speech or sound effects, and outputs these elements through a display or output unit, enhancing user immersion.
The solution enhances user immersion by providing artificial voices with intonation and emotion, improving understanding of the content, and accommodating users with disabilities through additional sensory cues.
Smart Images

Figure 0007714731000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a data processing apparatus, a data processing method, and a data processing program.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the prior art, there is still room for improvement in the quality of artificial voices output from generation models such as the above language model.
Means for Solving the Problems
[0005] The data processing apparatus according to the first aspect includes a processor. The processor acquires electronic content in which content including text and illustrations is digitized. A prompt including an image of a specific section among those partitioned in a predetermined section unit in the content and an instruction to estimate the emotion of a character shown in the image of the specific section is input to a generation model that generates information according to input data. When the image of the specific section is displayed on a display unit capable of displaying the electronic content, a speech bubble of the character in the specific section is output from an output unit as artificial speech generated based on the emotion of the character estimated by the generation model.
[0006] In the data processing apparatus according to the first aspect, the processor acquires electronic content. A prompt including an image of a specific section in the electronic content and an instruction to estimate the emotion of a character shown in the image of the specific section is input to the generation model. Then, when the image of the specific section is displayed on the display unit, a speech bubble of the character in the specific section is output from the output unit as artificial speech generated based on the emotion of the character estimated by the generation model. Thereby, according to this data processing apparatus, compared with a configuration in which artificial speech without intonation is output from the output unit, the immersion of the user in the electronic content can be enhanced.
[0007] The data processing apparatus according to the second aspect, in the first aspect, when onomatopoeia is included in the image of the specific section, the processor inputs a prompt including the image of the specific section and an instruction to interpret the onomatopoeia to the generation model. When the image of the specific section is displayed on the display unit, a sound effect generated based on the interpretation result of the onomatopoeia output by the generation model is output from the output unit.
[0008] In the data processing apparatus according to the second aspect, when an onomatopoeia is included in the image of a specific section, a prompt including the image of the specific section and an instruction to interpret the onomatopoeia is input to the generation model. Then, when the image of the specific section is displayed on the display unit, a sound effect generated based on the interpretation result of the onomatopoeia output by the generation model is output from the output unit. Thereby, according to the data processing apparatus, compared with a configuration in which a sound effect corresponding to the onomatopoeia is not output from the output unit, the immersion of the user in the electronic content can be enhanced.
[0009] In the data processing apparatus according to the third aspect, in the first aspect or the second aspect, when the processor receives a predetermined operation by the user while the image of the specific section is being displayed on the display unit, the processor causes the output unit to output, in a predetermined artificial voice, the estimated content of the emotion of the character generated by the generation model.
[0010] In the data processing apparatus according to the third aspect, when a predetermined operation by the user is received while the image of the specific section is being displayed on the display unit, the estimated content of the emotion of the character generated by the generation model is output from the output unit in a predetermined voice. Thereby, according to the data processing apparatus, compared with a configuration in which only the lines of the character are output as voice, the degree of understanding of the user with respect to the content of the electronic content can be enhanced.
[0011] In the data processing apparatus according to the fourth aspect, in any one of the first aspect to the third aspect, when the output of the sound corresponding to the image of the specific section from the output unit is completed, the processor performs at least one of outputting a specific sound and generating a specific vibration by using at least one of the sound output function of the output unit and the vibration function of the vibration unit.
[0012] In the data processing apparatus according to the fourth aspect, when the output of the sound corresponding to the image of the specific section from the output unit is completed, at least one of the sound output function and the vibration function is used to perform at least one of the output of the specific sound and the generation of the specific vibration. Thereby, according to the data processing apparatus, even if the user has a disability in vision, it is possible to make the user grasp that the switching timing of the section to be displayed on the display unit has arrived.
[0013] In the data processing apparatus according to the fifth aspect, in any one of the first aspect to the fourth aspect, when the content is visualized, the processor extracts a specific virtual actor corresponding to a specific actor in charge of the voice of the character from an actor database in which a plurality of virtual actors capable of outputting artificial voices with a predetermined voice quality are stored, and when the image of the specific section is displayed on the display unit, the processor causes the output unit to output the lines of the character in the specific section as artificial voices by the specific virtual actor.
[0014] In the data processing apparatus according to the fifth aspect, when the content is visualized, a specific virtual actor corresponding to a specific actor in charge of the voice of the character is extracted from the actor database. And when the image of the specific section is displayed on the display unit, the lines of the character in the specific section are output from the output unit as artificial voices by the specific virtual actor. Thereby, according to the data processing apparatus, it is possible to reduce the sense of discomfort that a user who knows the voice of the character when visualized feels about the artificial voice of the character.
[0015] In the data processing apparatus according to the sixth aspect, in the fifth aspect, when there are a plurality of the specific actors, the processor accepts a user's selection of one actor from among the plurality of the specific actors, extracts a first virtual actor corresponding to the one actor that has been selected from the actor database, and when the image of the specific section is displayed on the display unit, causes the output unit to output the lines of the character in the specific section as artificial voices by the first virtual actor.
[0016] In the data processing device according to the sixth aspect, when there are a plurality of specific actors, the user selects one actor from among the plurality of specific actors. From the actor database, a first virtual actor corresponding to the one actor selected by the user is extracted. Then, when an image of a specific section is displayed on the display unit, from the output unit, the lines of the character in the specific section are output as artificial voices by the first virtual actor. Thereby, according to the data processing device, when there are a plurality of specific actors, an artificial voice that matches the user's preference can be set for the character.
[0017] In the data processing device according to the seventh aspect, in any one of the first aspect to the sixth aspect, when the content is not visualized, the processor inputs the electronic content into the generation model and obtains the characteristics of the character interpreted thereby, and inputs a prompt including the obtained characteristics of the character and an instruction to inquire about an actor having a voice quality suitable for the characteristics into the generation model, extracts a second virtual actor corresponding to the actor output by the generation model from an actor database in which a plurality of virtual actors capable of outputting artificial voices of a predetermined voice quality are stored, and when an image of the specific section is displayed on the display unit, causes the output unit to output the lines of the character in the specific section as artificial voices by the second virtual actor.
[0018] In the data processing device according to the seventh aspect, when the content is not visualized, the characteristics of the character interpreted by inputting the electronic content into the generation model are obtained. A prompt including the characteristics of the character and an instruction to inquire about an actor having a voice quality suitable for the characteristics is input into the generation model. From the actor database, a virtual actor corresponding to the actor output by the generation model is extracted. Then, when an image of a specific section is displayed on the display unit, from the output unit, the lines of the character in the specific section are output as artificial voices by the virtual actor. Thereby, according to the data processing device, even when the content is not visualized, an artificial voice capable of reproducing a voice quality suitable for the characteristics of the character can be set for the character.
[0019] The data processing method of the eighth aspect acquires electronic content in which content including text and illustrations is digitized, and inputs, into a generation model that generates information according to input data, a prompt including an image of a specific section among those partitioned in a predetermined section unit in the content and an instruction to estimate the emotion of a character shown in the image of the specific section. When the image of the specific section is displayed on a display unit capable of displaying the electronic content, a computer executes a process of outputting, from an output unit, the lines of the character in the specific section as artificial voices generated based on the emotion of the character estimated by the generation model.
[0020] In the data processing method of the eighth aspect, electronic content is acquired. A prompt including an image of a specific section in the electronic content and an instruction to estimate the emotion of a character shown in the image of the specific section is input into the generation model. Then, when the image of the specific section is displayed on the display unit, the lines of the character in the specific section are output from the output unit as artificial voices generated based on the emotion of the character estimated by the generation model. Thereby, according to this data processing method, compared with a configuration in which artificial voices without intonation are output from the output unit, the sense of immersion of the user in the electronic content can be enhanced.
[0021] The data processing program of the ninth aspect causes a computer to execute a process of acquiring electronic content in which content including text and illustrations is digitized, inputting, into a generation model that generates information according to input data, a prompt including an image of a specific section among those partitioned in a predetermined section unit in the content and an instruction to estimate the emotion of a character shown in the image of the specific section, and outputting, from an output unit, the lines of the character in the specific section as artificial voices generated based on the emotion of the character estimated by the generation model when the image of the specific section is displayed on a display unit capable of displaying the electronic content.
[0022] In the data processing program of the ninth aspect, electronic content is acquired. A prompt including an image of a specific section in the electronic content and an instruction to estimate the emotion of a character shown in the image of the specific section is input to the generation model. When an image of the specific section is displayed on the display unit, a line of dialogue of the character in the specific section is output from the output unit as artificial speech generated based on the emotion of the character estimated by the generation model. Thereby, according to this data processing program, compared with a configuration in which artificial speech without intonation is output from the output unit, the sense of immersion of the user in the electronic content can be enhanced.
Brief Description of Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Best Mode for Carrying Out the Invention
[0024] Hereinafter, an example of an embodiment of a data processing apparatus, a data processing method, and a program according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0025] First, the terms used in the following description will be explained.
[0026] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), or an APU (Accelerated Processing Unit).
[0027] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0028] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0029] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to an embodiment.
[0032] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server. An example of the smart device 14 is a smartphone. In the present embodiment, the data processing device 12 is an example of the "data processing device" according to the technology of the present disclosure.
[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28 is an example of a "processor" according to the technology of the present disclosure. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.
[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by contact of an indicator (e.g., a pen or a finger, etc.) by detecting the contact of the indicator. The microphone 38B receives user input by voice by detecting the voice of the user. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires data indicating the user input.
[0036] The output device 40 includes a display 40A, a speaker 40B, etc., and presents data to the user by outputting the data in a form that can be perceived by the user (for example, voice and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The display 40A is an example of the "display unit" according to the technology of the present disclosure, and the speaker 40B is an example of the "output unit" according to the technology of the present disclosure.
[0037] The camera 42 is a small digital camera equipped with an optical system such as a lens, a diaphragm, and a shutter, and an imaging device such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0038] The communication I / F 44 is connected to the network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0040] As shown in FIG. 2, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of the "data processing program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as the specific processing unit 290 according to the specific processing program 56 executed on the RAM 30. Note that the "data processing program" according to the technology of the present disclosure can also be applied as a program product.
[0041] The data generation model 58 is stored in the storage 32. The data generation model 58 is used by the specific processing unit 290.
[0042] The data generation model 58 is a so-called generative AI (Artificial Intelligence). As an example of the data generation model 58, ChatGPT (Internet search <URL: https: / / openai.com / blog / chatgpt>), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) Examples of generative AI include EmotiVoice (Internet search <URL: https: / / weel.co.jp / media / tech / emotivoice / >), and Audiobox (Internet search <URL: https: / / audiobox.metademolab.com / >). The data generation model 58 is configured by appropriately combining various known generative AIs as described above. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is input thereto. The data generation model 58 infers the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The data generation model 58 is an example of the "generation model" according to the technology of the present disclosure.
[0043] In the smart device 14, reception / output processing is performed by the processor 46. A reception / output program 60 is stored in the storage 50. The reception / output program 60 is used in combination with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception / output program 60 from the storage 50 and executes the read reception / output program 60 on the RAM 48. The reception / output processing is realized by operating as the control unit 46A according to the reception / output program 60 executed by the processor 46 on the RAM 48.
[0044] Next, an example of specific processing performed by the data processing device 12 will be described. FIG. 3 is a first explanatory diagram showing an example of the identification process performed by the data processing device 12. FIG. 3 shows an example in which, in the identification process, a data generation model 58 estimates the emotion of a character shown in an image of a specific frame (hereinafter simply referred to as a "frame") of an electronic comic. The data generation model 58 is, for example, ChatGPT. An electronic comic is a paper comic (hereinafter simply referred to as a "comic") that includes text and illustrations and has been digitized. In the comic, each page is divided into predetermined frame units. An electronic comic is an example of "electronic content" according to the technology of the present disclosure, a comic is an example of "content" according to the technology of the present disclosure, and a frame unit is an example of a "single section unit" according to the technology of the present disclosure.
[0045] 3 shows a prompt 70 to be input to the data generation model 58. The prompt 70 includes a frame image 70A showing an image of a specific frame in the comic A as the comic, and an instruction statement 70B showing an instruction to estimate the emotion of the character shown in the frame image 70A. The frame image 70A is an example of an "image of a specific section" according to the technology of the present disclosure.
[0046] Frame image 70A shows a character saying, "I wonder if it will be sunny tomorrow." In addition, instruction sentence 70B is, for example, text that reads, "How do you think the character in this frame is feeling?"
[0047] 3 also shows an output result 71 when a prompt 70 is input to the data generation model 58. An example of the output result 71 is the text "The image shows a teary-eyed character standing still. From this, it can be inferred that the character is feeling anxious or sad."
[0048] Fig. 4 is a second explanatory diagram showing an example of the identification process performed by the data processing device 12. Fig. 4 shows an example in which the data generation model 58 interprets onomatopoeia shown in an image of a specific frame of a digital comic in the identification process. The data generation model 58 is, for example, ChatGPT.
[0049] Figure 4 shows a prompt 72 to be input into the data generation model 58. The prompt 72 is composed of a frame image 72A showing an image of a specific frame in Comic A and an instruction text 72B showing an instruction to interpret the onomatopoeia shown in the frame image 72A. The frame image 72A is an example of the "image of a specific section" according to the technology of the present disclosure.
[0050] The frame image 72A shows an onomatopoeia of "Dogyun". Also, the instruction text 72B is, as an example, the text of "What sound is the onomatopoeia of this frame?".
[0051] Also, Figure 4 shows an output result 73 when the prompt 72 is input into the data generation model 58. The output result 73 is, as an example, the text of "The onomatopoeia 'Dogyun' is an onomatopoeic word indicating a large impact or an object moving at high speed, etc.".
[0052] Next, the operation of the data processing system 10 will be described. An example of the flow of a specific process will be described with reference to FIGS. 5 to 10. The specific process includes a first specific process shown in FIGS. 5 to 9 and a second specific process shown in FIG. 10. As an example, the first specific process is performed when the user operates the smart device 14 to execute a predetermined application installed in the smart device 14 and a selection screen for selecting an electronic comic is displayed on the display 40A. Also, the second specific process is performed when the user operates the smart device 14 to execute the predetermined application and a browsing screen for browsing the electronic comic is displayed on the display 40A. Note that the flow of the specific process shown in FIGS. 5 to 10 is an example of the "data processing method" according to the technology of the present disclosure.
[0053] In step S10 shown in FIG. 5, the processor 28 obtains a specific electronic comic from the database 24 based on user input for the smart device 14, for example, data indicating user input by voice. The database 24 stores various electronic comics in which various comics are digitized. Hereinafter, the specific electronic comic obtained from the database 24 will be described as "Electronic Comic A" in which Comic A is digitized. Then, the processor 28 proceeds to step S11.
[0054] In step S11, the processor 28 lists the characters appearing in Electronic Comic A. Here, the processor 28 inputs a prompt including Electronic Comic A, an instruction to interpret the story and style of Electronic Comic A, and an instruction to list the characteristics of each character appearing in Electronic Comic A into the data generation model 58, and obtains the output result. The data generation model 58 is, for example, ChatGPT. The prompt is, for example, text and PDF data such as "Please interpret the story and style of the following Electronic Comic A. Also, please list the characteristics of each character appearing in Electronic Comic A. Electronic Comic A.pdf". As a result, in step S11, based on the output result by the data generation model 58, the characteristics of each character such as "Character A: a boy about 10 years old with a quiet personality, Character B: a boy about 10 years old with an active personality..." are listed. Then, the processor 28 proceeds to step S12.
[0055] In step S12, the processor 28 determines the BGM (background music) output from the speaker 40B while the electronic comic A is being displayed on the display 40A. The database 24 stores various BGMs corresponding to the styles of various electronic comics. The processor 28 obtains from the database 24 the BGM corresponding to the style of the electronic comic A shown in the output result of the data generation model 58 in step S11. The processor 28 transmits the sound data indicating the obtained BGM to the processor 46. Then, the processor 28 proceeds to step S13.
[0056] In step S13, the processor 28 performs a frame information generation process for generating frame information to be included in each frame of the electronic comic A. The frame information includes various types of information described later in addition to the information indicating the frame image. The subroutine of the frame information generation process will be described later. Then, the processor 28 proceeds to step S14.
[0057] In step S14, the processor 28 performs an emotion estimation process for estimating the emotion of each character appearing in the electronic comic A in each frame. The subroutine of the emotion estimation process will be described later. Then, the processor 28 proceeds to step S15.
[0058] In step S15, the processor 28 performs a sound effect generation process for generating a sound effect corresponding to the onomatopoeia shown in a specific frame of the electronic comic A. The subroutine of the sound effect generation process will be described later. Then, the processor 28 proceeds to step S16.
[0059] In step S16, the processor 28 performs a synthetic voice determination process for determining the synthetic voice that utters the lines of the characters appearing in the electronic comic A. The subroutine of the synthetic voice determination process will be described later. Then, the processor 28 ends the process.
[0060] Figure 6 is a subroutine of the frame information generation process. In step S20 shown in FIG. 6, the processor 28 divides the electronic comic A into each frame. Then, the processor 28 proceeds to step S21. As an example, the order of each frame of the electronic comic A is predetermined, and the processor 28 performs the processing after step S21 in the predetermined order of the frames.
[0061] In step S21, the processor 28 identifies the characters appearing in the frame. For example, the processor 28 refers to the image of the frame and the characteristics of each character listed in step S11 to identify the characters appearing in the frame. Then, the processor 28 proceeds to step S22.
[0062] In step S22, the processor 28 performs character recognition processing on the image of the frame to identify the lines of the characters appearing in the frame. When there are multiple characters appearing in the frame image and multiple lines, the processor 28 identifies the speaker of each line based on the image information and character recognition information, etc. Then, the processor 28 proceeds to step S23.
[0063] In step S23, the processor 28 associates the characters identified in step S21 and the lines of the characters identified in step S22 with the frame information of the corresponding frame and stores them in the database 24. Thereby, in the database 24, for example, in the first frame of the electronic comic A, information that character A appears and the line of character A is "Will it be sunny tomorrow?" is stored. Then, the processor 28 proceeds to step S24.
[0064] In step S24, the processor 28 determines whether the frame information stored in the database 24 in step S23 corresponds to the last frame of the electronic comic A. Here, when the processor 28 determines that the frame information corresponds to the last frame of the electronic comic A (step S24: YES), it returns to the calling process. On the other hand, when the processor 28 determines that the frame information does not correspond to the last frame of the electronic comic A (step S24: NO), it proceeds to step S25. As an example, a specific label is assigned to the last frame of the electronic comic A, and the processor 28 determines whether it corresponds to the last frame based on the presence or absence of the specific label.
[0065] In step S25, the processor 28 advances the processing target to the next frame. Then, the processor 28 returns to step S21. In this way, the processor 28 repeatedly executes the subroutine shown in FIG. 6 from the first frame to the last frame of the electronic comic A.
[0066] FIG. 7 is a subroutine for emotion estimation processing. In step S30 shown in FIG. 7, the processor 28 acquires the frame information corresponding to the first frame of the electronic comic A from the database 24. As an example, a predetermined label is assigned to the first frame of the electronic comic A, and the processor 28 acquires the frame information corresponding to the frame with the predetermined label from the database 24. Then, the processor 28 proceeds to step S31.
[0067] In step S31, the processor 28 generates a prompt to be input to the data generation model 58. The data generation model 58 is, for example, ChatGPT. The prompt includes an image of a frame to be processed of the electronic comic A and an instruction to estimate the emotion of the character shown in the image. The frame to be processed is the first frame in the first time of the subroutine shown in FIG. 7, and is the frame corresponding to the frame information acquired in the steps within the process from the second time onward. The "frame to be processed" that appears hereinafter has the same meaning. For example, the prompt is configured to include the image shown in the frame image 70A of FIG. 3 and the text shown in the instruction text 70B. Then, the processor 28 proceeds to step S32.
[0068] In step S32, the processor 28 inputs the prompt generated in step S31 to the data generation model 58 and obtains an output result by the data generation model 58. Then, the processor 28 proceeds to step S33.
[0069] In step S33, the processor 28 associates the output result output from the data generation model 58 in step S32 with the frame information of the frame to be processed and stores it in the database 24. For example, the output result is text as shown in the output result 71 of FIG. 3. Then, the processor 28 proceeds to step S34.
[0070] In step S34, the processor 28 determines whether the frame information stored in the database 24 in step S33 corresponds to the last frame of the electronic comic A. Here, when the processor 28 determines that the frame information corresponds to the last frame of the electronic comic A (step S34: YES), it returns to the calling process. On the other hand, when the processor 28 determines that the frame information does not correspond to the last frame of the electronic comic A (step S34: NO), it proceeds to step S35.
[0071] In step S35, the processor 28 advances the processing target to the next frame and acquires the frame information corresponding to the next frame from the database 24. Then, the processor 28 returns to step S31. In this way, the processor 28 repeatedly executes the subroutine shown in FIG. 7 from the first frame to the last frame of the electronic comic A.
[0072] FIG. 8 is a subroutine for sound effect generation processing. In step S40 shown in FIG. 8, the processor 28 acquires the frame information corresponding to the first frame of the electronic comic A from the database 24. Then, the processor 28 proceeds to step S41.
[0073] In step S41, the processor 28 determines whether or not the frame being processed contains onomatopoeia. Here, when the processor 28 determines that onomatopoeia is included (step S41: YES), it proceeds to step S42. On the other hand, when the processor 28 determines that onomatopoeia is not included (step S41: NO), it proceeds to step S47. The database 24 stores various onomatopoeia used in various comics. The processor 28 performs character recognition processing on the image of the frame being processed and determines the presence or absence of onomatopoeia based on whether the result of the character recognition processing matches the onomatopoeia stored in the database 24.
[0074] In step S42, the processor 28 generates a prompt to be input to the data generation model 58. The data generation model 58 is, for example, ChatGPT. The prompt includes the image of the frame being processed and an instruction to interpret the onomatopoeia shown in the image. For example, the prompt is configured to include the image shown in the frame image 72A of FIG. 4 and the text shown in the instruction text 72B. Then, the processor 28 proceeds to step S43.
[0075] In step S43, the processor 28 inputs the prompt generated in step S42 into the data generation model 58 and obtains the output result by the data generation model 58. Then, the processor 28 proceeds to step S44.
[0076] In step S44, the processor 28 generates a prompt to be input into the data generation model 58. The data generation model 58 is, for example, Audiobox. The prompt includes the interpretation result of the onomatopoeia, which is the output result of the data generation model 58 in step S43, and an instruction to generate a sound effect corresponding to the interpretation result of the onomatopoeia. For example, the interpretation result of the onomatopoeia is text as shown in the output result 73 of FIG. 4. As a result, the prompt becomes text such as "The onomatopoeia 'Dogyun' is an onomatopoeic word indicating a large impact or an object moving at high speed. Please generate a sound effect suitable for this onomatopoeic word." Then, the processor 28 proceeds to step S45.
[0077] In step S45, the processor 28 inputs the prompt generated in step S44 into the data generation model 58 and obtains the output result by the data generation model 58. Then, the processor 28 proceeds to step S46.
[0078] In step S46, the processor 28 associates the output result output from the data generation model 58 in step S45 with the frame information of the frame to be processed and stores it in the database 24. For example, the output result is sound data indicating a sound effect. Then, the processor 28 proceeds to step S47.
[0079] In step S47, the processor 28 determines whether the frame information stored in the database 24 in step S46 corresponds to the last frame of the electronic comic A. Here, when the processor 28 determines that the frame information corresponds to the last frame of the electronic comic A (step S47: YES), it returns to the calling process. On the other hand, when the processor 28 determines that the frame information does not correspond to the last frame of the electronic comic A (step S47: NO), it proceeds to step S48.
[0080] In step S48, the processor 28 advances the processing target to the next frame and acquires the frame information corresponding to the next frame from the database 24. Then, the processor 28 returns to step S41. In this way, the processor 28 repeatedly executes the subroutine shown in FIG. 7 from the first frame to the last frame of the electronic comic A.
[0081] FIG. 9 is a subroutine of the artificial voice determination process. In step S50 shown in FIG. 9, the processor 28 selects an arbitrary character from the characters appearing in the electronic comic A. Then, the processor 28 proceeds to step S51.
[0082] In step S51, the processor 28 determines whether the comic A has been animated. Here, when the processor 28 determines that the comic A has been animated (step S51: YES), it proceeds to step S52. On the other hand, when the processor 28 determines that the comic A has not been animated (step S51: NO), it proceeds to step S54. As an example, the processor 28 determines whether the comic A has been animated based on the output result of the data generation model 58. The data generation model 58 is, for example, ChatGPT. In this case, the processor 28 inputs a prompt asking whether the comic A has been animated, such as "Has Comic A been animated?", to the data generation model 58 and obtains the output result by the data generation model 58. Animation is an example of "video conversion" according to the technology of the present disclosure.
[0083] In step S52, the processor 28 identifies the voice actor who dubbed the voice of the character to be processed. The character to be processed is the character selected in step S50 in the first time of the subroutine shown in FIG. 9, and the character selected in step S59 in the second time and later. The "character to be processed" that appears later also has the same meaning. As an example, the processor 28 identifies the voice actor based on the output result of the data generation model 58. The data generation model 58 is, for example, ChatGPT. In this case, the processor 28 inputs a prompt asking the data generation model 58 "Who is the voice actor who dubbed the voice of the character to be processed in the anime of Comic A?" and obtains the output result by the data generation model 58. Then, the processor 28 proceeds to step S53.
[0084] In step S53, the processor 28 determines a virtual voice actor who utters the lines of the character to be processed. The database 24 stores a voice actor database in which a plurality of virtual voice actors capable of outputting artificial voices of a predetermined voice quality are stored. A virtual voice actor is a virtual voice actor trained to be able to utter an artificial voice of the same type as a real voice actor based on the voice of the real voice actor. The predetermined voice quality is, for example, a clear and firm voice, a high-pitched childlike voice, a lively and bright voice, a gentle and cute voice, a low and cool voice, a refreshing young man's voice, a boy's voice immediately after voice change, a deep and bass voice, a soft and warm voice, and an elegant adult voice, etc. The voice actor database stores information such as the voice actor corresponding to the virtual voice actor (in other words, the voice actor on which the artificial voice of the virtual voice actor is based), the characters for which the voice actor has dubbed, and the voice quality corresponding to the virtual voice actor, associated with each virtual voice actor. Thus, the voice actor database stores information such as "the voice actor corresponding to the virtual voice actor A is the voice actor A, the characters for which the voice actor A has dubbed are the character A, the character C, and the character E, etc., and the voice quality corresponding to the virtual voice actor A is a clear and firm voice".
[0085] Here, the processor 28 extracts a specific virtual voice actor corresponding to the voice actor identified in step S52 from the voice actor database. As a result, when the processor 28 identifies voice actor A in step S52, it determines that the virtual voice actor A corresponding to voice actor A will be the virtual voice actor that utters the lines of the character to be processed. Then, the processor 28 proceeds to step S57. The voice actor database is an example of the "actor database" according to the technology of the present disclosure, the voice actor is an example of the "actor" according to the technology of the present disclosure, the virtual voice actor is an example of the "virtual actor" according to the technology of the present disclosure, and the virtual voice actor A is an example of the "specific virtual actor" according to the technology of the present disclosure.
[0086] In step S54, the processor 28 generates a prompt to be input to the data generation model 58. The data generation model 58 is, for example, ChatGPT. The prompt includes the characteristics of the character to be processed and an instruction to inquire about a voice actor having a voice quality suitable for the characteristics. The characteristics of the character to be processed use the information listed in step S11 shown in FIG. 5. As a result, the prompt becomes text such as "Character B is a boy of about 10 years old and has an active personality. Who is the voice actor having a voice quality suitable for the characteristics of this character B?" Then, the processor 28 proceeds to step S55.
[0087] In step S55, the processor 28 inputs the prompt generated in step S54 to the data generation model 58 and obtains the output result by the data generation model 58. Then, the processor 28 proceeds to step S56.
[0088] In step S56, the processor 28 determines a virtual voice actor who utters the lines of the character to be processed. Here, the processor 28 extracts from the voice actor database a virtual voice actor corresponding to the voice actor shown in the output result of the data generation model 58 in step S55. Thereby, when the output result of the data generation model 58 is voice actor B, the processor 28 determines the virtual voice actor B corresponding to voice actor B as the virtual voice actor who utters the lines of the character to be processed. Then, the processor 28 proceeds to step S57. The virtual voice actor B is an example of the "second virtual performer" according to the technology of the present disclosure.
[0089] In step S57, the processor 28 associates the character to be processed with the virtual voice actor who utters the lines of the character and stores the association in the database 24. Then, the processor 28 proceeds to step S58.
[0090] In step S58, the processor 28 determines whether the association with the virtual voice actor for all characters has been completed. Here, when the processor 28 determines that the association with the virtual voice actor for all characters has been completed (step S58: YES), it returns to the calling process. On the other hand, when the processor 28 determines that the association with the virtual voice actor for all characters has not been completed (step S58: NO), it proceeds to step S59.
[0091] In step S59, the processor 28 selects the next character to be processed. Then, the processor 28 returns to step S51. In this way, the processor 28 repeatedly executes the subroutine shown in FIG. 9 until the association with the virtual voice actor for all characters appearing in the electronic comic A is completed.
[0092] FIG. 10 is a flowchart showing the flow of the second specific process. In step S60 shown in FIG. 10, the processor 28 acquires from the database 24 the frame information corresponding to the frame specified by the user. Then, the processor 28 proceeds to step S61.
[0093] In step S61, the processor 28 transmits the frame information corresponding to the frame specified by the user in step S60 to the smart device 14. As an example, the processor 28 transmits at least the image of the frame, the character recognition result of the lines of the characters in the frame, and the estimated result of the emotions of the characters appearing in the frame among the frame information, and when onomatopoeia is included in the frame, additionally transmits sound data indicating the sound effect corresponding to the onomatopoeia. Thereby, the image of the frame is displayed on the display 40A of the smart device 14. Further, based on the display of the image of the frame, the processor 46 of the smart device 14 outputs the BGM indicated in the acquired sound data from the speaker 40B. Then, the processor 28 proceeds to step S62.
[0094] In step S62, the processor 28 instructs the smart device 14 to have the virtual voice actor that utters the lines of the character appearing in the current frame based on the frame information acquired in step S60. The processor 28 determines the virtual voice actor that utters the lines of the current frame based on the frame information and the association between each character and each virtual voice actor stored in the database 24, and notifies the determined virtual voice actor to the processor 46. In the storage 50 of the smart device 14, there is stored text-to-speech software that can read text in the artificial voice of each virtual voice actor registered in the voice actor database. The processor 46 of the smart device 14 sets the virtual voice actor that utters the lines of the current frame from among the text-to-speech software according to the instruction from the processor 28. Then, the processor 46 uses the text-to-speech software to generate sound data in which the set virtual voice actor reads the text in an artificial voice based on the estimated result of the emotion of the character included in the frame information transmitted from the processor 28, and instructs the output of the sound data to the speaker 40B. As a result, from the speaker 40B, the lines of the character in the frame are output in the artificial voice of the set virtual voice actor. Also, since the virtual voice actor reads the text based on the estimated result of the emotion of the character included in the frame information transmitted from the processor 28, from the speaker 40B, the lines of the character are output in an artificial voice that reflects the emotion of the character estimated by the data generation model 58. Also, when the frame includes onomatopoeia, from the speaker 40B, a sound effect corresponding to the onomatopoeia shown in the acquired sound data is output. Although detailed description here is omitted, when the output of the artificial voice indicating the lines of the character in the frame or the sound effect corresponding to the onomatopoeia ends, the processor 46 of the smart device 14 outputs a specific sound similar to step S66 described later from the speaker 40B. Then, the processor 28 proceeds to step S63.
[0095] In step S63, the processor 28 determines whether there is a change in the frames of the electronic comic A to be displayed on the display 40A. Here, when the processor 28 determines that there is a change in the frames (step S63: YES), it proceeds to step S67. On the other hand, when the processor 28 determines that there is no change in the frames (step S63: NO), it proceeds to step S64. As an example, when the processor 28 acquires data indicating user input by a flick operation on the display 40A, it determines that there is a change in the frames. The data indicating user input on the display 40A is appropriately transmitted from the smart device 14 to the data processing device 12.
[0096] In step S64, the processor 28 determines whether a predetermined operation has been performed on the display 40A. Here, when the processor 28 determines that a predetermined operation has been performed (step S64: YES), it proceeds to step S65. On the other hand, when the processor 28 determines that a predetermined operation has not been performed (step S64: NO), it proceeds to step S68. As an example, when the processor 28 acquires data indicating user input by a tap operation on the display 40A, it determines that a predetermined operation has been performed.
[0097] In step S65, the processor 28 instructs the smart device 14 to output the explanation of the frame to be processed. The explanation of the frame includes the estimated result of the emotion of the character included in the frame information and the estimated content of the situation of the character based on the analysis result of the image of the frame. As a result, from the speaker 40B, the explanation of the frame is output in a predetermined artificial voice. The explanation of the frame includes, for example, the text content as shown in the output result 71 of FIG. 3 as the estimated result of the emotion of the character. The predetermined artificial voice may be an artificial voice by any virtual voice actor in the text reading software, or an artificial voice imitating the voice of a person (e.g., mother, father) specified by the user in advance. When the predetermined artificial voice is an artificial voice imitating the voice of a person specified by the user in advance, the smart device 14 is previously set to be able to output the artificial voice from the speaker 40B. Then, the processor 28 proceeds to step S66.
[0098] In step S66, the processor 28 instructs the smart device 14 to output a specific sound. The specific sound is a sound corresponding to the image of the frame to be processed, for example, a sound for notifying the user that the output of the artificial voice indicating the explanation of the frame, the artificial voice indicating the lines of the character in the frame, and the sound effect corresponding to the onomatopoeia of the frame has ended. The type of the specific sound is not particularly limited. As a result, a specific sound is output from the speaker 40B. Then, the processor 28 returns to step S63. Note that the data indicating the presence or absence of the sound output from the speaker 40B is appropriately transmitted from the smart device 14 to the data processing device 12.
[0099] In step S67, the processor 28 advances the frame to be processed to the next frame. Then, the processor 28 returns to step S60.
[0100] In step S68, the processor 28 determines whether the end condition of a predetermined application is satisfied. Here, when the processor 28 determines that the end condition is satisfied (step S68: YES), the process ends. On the other hand, when the processor 28 determines that the end condition is not satisfied (step S68: NO), it returns to step S63. As an example, when the processor 28 acquires data indicating a user input corresponding to an end operation for ending a predetermined application, it determines that the end condition is satisfied.
[0101] Next, an example of data output from the output device 40 of the smart device 14 based on the execution of specific processing will be described.
[0102] FIG. 11 is a first explanatory diagram showing a display example of the display 40A. On the display 40A shown in FIG. 11, a frame image 80 showing an image of the first frame of the electronic comic A is displayed. The frame image 80 includes a character C1 and a line 80A of the character C1. The character C1 is a human character. The content of the line 80A is "I wonder if it will be sunny tomorrow."
[0103] At this time, based on the fact that the frame image 80 is displayed on the display 40A, from the speaker 40B, the content of the text shown in the line 80A is output as artificial voice of a virtual voice actor (e.g., virtual voice actor E) corresponding to the character C1. Further, since the virtual voice actor E reads the text based on the estimated result of the emotion of the character C1, artificial voice reflecting the emotion of the character C1 is output from the speaker 40B.
[0104] FIG. 12 is a second explanatory diagram showing a display example of the display 40A. On the display 40A shown in FIG. 12, a frame image 81 showing an image of the second frame of the electronic comic A is displayed. The frame image 81 includes an onomatopoeia 81A. The onomatopoeia 81A is an onomatopoeic word such as "Dogyun".
[0105] At this time, based on the fact that the frame image 81 is displayed on the display 40A, a sound effect generated by the data generation model 58 corresponding to the onomatopoeia 81A is output from the speaker 40B. The data generation model 58 is, for example, Audiobox.
[0106] FIG. 13 is a third explanatory diagram showing a display example of the display 40A. A frame image 82 showing the third frame image of the electronic comic A is displayed on the display 40A shown in FIG. 13. The frame image 82 includes a character C1, a character C2, an onomatopoeia 82A, a line 82B of the character C2, and a line 82C of the character C1. The character C2 is a human character. The onomatopoeia 82A is an onomatopoeia for "jaan". The content of the line 82B is "It will surely clear up!" The content of the line 82C is "That's right, thank you!"
[0107] At this time, based on the fact that the frame image 82 is displayed on the display 40A, sounds corresponding to the frame image 82 are sequentially output from the speaker 40B in a predetermined order. In the present embodiment, when there are a plurality of sounds corresponding to the frame image, the sounds are sequentially output from the speaker 40B in a predetermined output order. As an example, in the third frame image, the corresponding sounds are sequentially output in the order of the onomatopoeia 82A, the line 82B, and the line 82C.
[0108] Based on the display of the frame image 82 on the display 40A, first, from the speaker 40B, the sound effect generated by the data generation model 58 corresponding to the onomatopoeia 82A is output. The data generation model 58 is, for example, Audiobox. Next, from the speaker 40B, the content of the text shown in the dialogue 82B is output in the artificial voice of the virtual voice actor (e.g., virtual voice actor B) corresponding to the character C2. Also, since the virtual voice actor B reads the text based on the estimated result of the emotion of the character C2, from the speaker 40B, an artificial voice reflecting the emotion of the character C2 is output. Finally, from the speaker 40B, the content of the text shown in the dialogue 82C is output in the artificial voice of the virtual voice actor (e.g., virtual voice actor E) corresponding to the character C1. Also, since the virtual voice actor E reads the text based on the estimated result of the emotion of the character C1, from the speaker 40B, an artificial voice reflecting the emotion of the character C1 is output.
[0109] As described above, in the data processing apparatus 12, the processor 28 acquires the electronic comic A in which the comic A is digitized. Also, the processor 28 inputs a prompt including the image of a specific frame among those partitioned in frame units predetermined in the comic A and an instruction to estimate the emotion of the character shown in the image of the specific frame to the data generation model 58. Then, when the image of a specific frame is displayed on the display 40A, the processor 28 causes the speaker 40B to output, in the artificial voice generated based on the emotion of the character estimated by the data generation model 58, the dialogue of the character in the specific frame. Thereby, according to the data processing apparatus 12, compared with a configuration in which a monotonous artificial voice is output from the speaker 40B, the immersion feeling of the user with respect to the electronic comic A can be enhanced. The specific frame is an example of the "specific partition" according to the technology of the present disclosure.
[0110] Also, in the data processing device 12, when the processor 28 detects that an onomatopoeia is included in the image of a specific frame, the processor 28 inputs a prompt including the image of the specific frame and an instruction to interpret the onomatopoeia to the data generation model 58. Then, when the image of the specific frame is displayed on the display 40A, the processor 28 causes the speaker 40B to output a sound effect generated based on the interpretation result of the onomatopoeia output by the data generation model 58. Accordingly, according to the data processing device 12, the immersion of the user in the electronic comic A can be enhanced as compared with a configuration in which a sound effect corresponding to the onomatopoeia is not output from the speaker 40B.
[0111] Also, in the data processing device 12, when the processor 28 receives a predetermined operation by the user while the image of a specific frame is being displayed on the display 40A, the processor 28 causes the speaker 40B to output an explanation of the frame including the estimated content of the character's emotion generated by the data generation model 58 in a predetermined synthetic voice. Accordingly, according to the data processing device 12, the user's understanding of the content of the electronic comic A can be enhanced as compared with a configuration in which only the character's lines are output as sound.
[0112] Also, in the data processing device 12, when the output of the sound corresponding to the image of a specific frame from the speaker 40B ends, the processor 28 uses the sound output function of the speaker 40B to output a specific sound. The sound corresponding to the image of the specific frame is at least one of a synthetic voice indicating the character's lines in the specific frame, a sound effect corresponding to the onomatopoeia in the specific frame, and a synthetic voice indicating the explanation in the specific frame. Accordingly, according to the data processing device 12, even if the user has a disability in vision, the user can be made aware that the timing of switching the frame to be displayed on the display 40A has arrived.
[0113] In the data processing device 12, when the processor 28 determines that the comic A has been animated, the processor 28 extracts a specific virtual voice actor corresponding to a specific voice actor who was responsible for the voice of the character from the voice actor database. Then, when an image of a specific frame is displayed on the display 40A, the processor 28 causes the speaker 40B to output the lines of the character in the specific frame as artificial voice by the specific virtual voice actor. Thereby, according to the data processing device 12, it is possible to reduce the sense of discomfort that a user who knows the voice of the character when it is animated feels about the artificial voice of the character.
[0114] In the data processing device 12, when the processor 28 determines that the comic A has not been animated, the processor 28 inputs the electronic comic A into the data generation model 58 to obtain the characteristics of the character interpreted. Further, the processor 28 inputs a prompt including the obtained characteristics of the character and an instruction to inquire about a voice actor having a voice quality suitable for the characteristics into the data generation model 58. Further, the processor 28 extracts a virtual voice actor corresponding to the voice actor output by the data generation model 58 from the voice actor database. Then, when an image of a specific frame is displayed on the display 40A, the processor 28 causes the speaker 40B to output the lines of the character in the specific frame as artificial voice by the virtual voice actor. Thereby, according to the data processing device 12, even if the comic A has not been animated, it is possible to set an artificial voice capable of reproducing a voice quality suitable for the characteristics of the character for the character. The virtual voice actor is an example of the "second virtual actor" according to the technology of the present disclosure.
[0115] (Others) When there are multiple specific voice actors who are in charge of the voices of the characters to be processed in the anime, the processor 28 may accept the selection of one voice actor by the user from among the multiple specific voice actors in the artificial voice determination process. In this case, the processor 28 extracts a virtual actor corresponding to the one voice actor who has received the selection by the user from the voice actor database. Then, when an image of a specific frame is displayed on the display 40A, the processor 28 causes the artificial voice by the virtual voice actor to be output from the speaker 40B for the lines of the character in the specific frame. Thereby, according to the data processing device 12, when there are multiple specific voice actors, an artificial voice that matches the user's preference can be set for the character. The virtual voice actor is an example of the "first virtual actor" according to the technology of the present disclosure.
[0116] In the above embodiment, the electronic comic is taken as an example of the "electronic content" according to the technology of the present disclosure, but it is not limited thereto. For example, an example of the "electronic content" may be another e-book in which a paper textbook or picture book containing text and illustrations is digitized. Also, not limited to books, specific processing according to this embodiment may be made possible for product packages (e.g., snack bags).
[0117] In the above embodiment, the animation is taken as an example of the "visualization" according to the technology of the present disclosure, but it is not limited thereto. For example, an example of the "visualization" may be a movie adaptation or the like.
[0118] In the above embodiment, the voice actor is taken as an example of the "actor" according to the technology of the present disclosure, but it is not limited thereto. For example, an example of the "actor" may be an actor, actress, or idol, etc.
[0119] In the above embodiment, when the output of the sound corresponding to the image of a specific frame from the speaker 40B is completed, the processor 28 outputs a specific sound using the sound output function of the speaker 40B. Instead of or in addition to this, when the output of the sound corresponding to the image of a specific frame from the speaker 40B is completed, the processor 28 may generate a specific vibration using the vibration function of a vibration unit (not shown) provided in the smart device 14. The vibration unit is a known vibration mechanism such as a motor and a pendulum mounted on various smartphones. In this case, the processor 28 instructs the smart device 14 to generate a specific vibration based on the fact that the output of the sound corresponding to the image of a specific frame from the speaker 40B has ended. Thereby, a specific vibration is generated from the vibration unit.
[0120] In the above embodiment, among the specific processes, the first specific process shown in FIGS. 5 to 9 was performed in advance by the user before viewing the electronic comic, but it is not limited to this. For example, the first specific process may be processed in real time in parallel with the second specific process while the user is viewing the electronic comic.
[0121] In the above embodiment, during the viewing of the electronic comic by the user, it may be possible to reselect the virtual voice actor who utters the lines of the character. The reselection can be performed by user input via the reception device 38. Thereby, when the artificial voice output from the speaker 40B is different from the user's image, the user can be made to select the virtual voice actor until it matches the user's image.
[0122] In the above embodiment, it may be possible to select the language output from the speaker 40B. The selection can be performed by user input via the reception device 38. Thereby, it becomes possible to output voice in a language different from the language of the text described in the electronic comic.
[0123] In the above embodiment, based on the character recognition result of the character's lines in the frame and the estimated content of the character's emotion generated by ChatGPT included in the data generation model 58, the text-to-speech software on the smart device 14 side generates sound data for the artificial voice to read the character's lines. However, the method for generating the sound data is not limited to this. For example, the character recognition result of the character's lines and the estimated content of the character's emotion may be input into EmotiVoice included in the data generation model 58, and the sound data may be generated as the output of the data generation model 58. In this way, when the data generation model 58 on the data processing device 12 side generates the sound data, the processor 28 transmits the generated sound data to the processor 46 before or during the user's viewing of the electronic comic. Then, when the image of the frame is displayed on the display 40A, the processor 46 instructs the speaker 40B to output the acquired sound data, and causes the artificial voice indicated by the sound data to be output from the speaker 40B.
[0124] In the above embodiment, a prompt including the onomatopoeia interpretation result, which is the output result of ChatGPT included in the data generation model 58, and an instruction to generate a sound effect corresponding to the onomatopoeia interpretation result is input into Audiobox included in the data generation model 58 to generate sound data indicating the sound effect. However, the method for generating the sound data is not limited to this. For example, the sound data is not limited to being generated by the data generation model 58, and may be generated using known software capable of generating predetermined sound effects.
[0125] As described above, the data processing system 10 according to the present disclosure has been mainly described in terms of the functions of the data processing apparatus 12. However, the data processing system 10 is not necessarily implemented on a server. The data processing system 10 may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program operating on a personal computer or as an application operating on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).
[0126] In the above-described embodiment, an example has been given in which specific processing is performed by the computer 22 of one data processing apparatus 12. However, the technology of the present disclosure is not limited to this, and distributed processing for specific processing by a plurality of computers including the computer 22, for example, the computer 22 and the computer 36 of the smart device 14, may be performed. In this case, an example of the "processor" according to the technology of the present disclosure is the processor 28 and the processor 46.
[0127] In the above-described embodiment, an example has been given and described in which the specific processing program 56 is stored in the storage 32. However, the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing apparatus 12. The processor 28 executes specific processing according to the specific processing program 56.
[0128] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing apparatus 12 via the network 54, and the specific processing program 56 may be downloaded and installed in the computer 22 in response to a request from the data processing apparatus 12.
[0129] Note that it is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32. A part of the specific processing program 56 may be stored instead.
[0130] As hardware resources for executing specific processing, various types of processors shown below can be used. As a processor, for example, a CPU, which is a general-purpose processor that functions as a hardware resource for executing specific processing by executing software, that is, a program, can be mentioned. Also, as a processor, for example, a dedicated electric circuit, which is a processor having a circuit configuration designed specifically for executing specific processing such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), can be mentioned. A memory is built in or connected to any of these processors, and any of these processors executes specific processing by using the memory.
[0131] The hardware resources for executing specific processing may be configured by one of these various types of processors, or may be configured by a combination of two or more processors of the same type or different types (for example, a combination of a plurality of FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resources for executing specific processing may be a single processor.
[0132] As an example of a configuration consisting of one processor, first, one processor is configured by a combination of one or more CPUs and software, and this processor functions as a hardware resource for executing a specific process. Second, as represented by a SoC (System-on-a-chip), there is a form in which a processor that realizes the functions of the entire system including a plurality of hardware resources for executing a specific process is used on one IC chip. Thus, the specific process is realized as a hardware resource using one or more of the above various processors.
[0133] Furthermore, as a hardware structure of these various processors, more specifically, an electric circuit combining circuit elements such as semiconductor elements can be used. Also, the above specific process is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be changed within the scope of not departing from the gist.
[0134] The above-described description and illustration are detailed descriptions of the part related to the technology of the present disclosure and are merely examples of the technology of the present disclosure. For example, the description regarding the above configuration, function, action, and effect is an example of the configuration, function, action, and effect of the part related to the technology of the present disclosure. Therefore, it goes without saying that within the scope of not departing from the gist of the technology of the present disclosure, the above-described description and illustration may be deleted of unnecessary parts, new elements may be added, or replacements may be made. Also, in order to avoid complication and facilitate the understanding of the part related to the technology of the present disclosure, the description regarding common technical knowledge etc. that does not particularly require explanation for enabling the implementation of the technology of the present disclosure is omitted in the above-described description and illustration.
[0135] All documents, patent applications, and technical standards described in this specification are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually stated to be incorporated by reference.
Explanation of Signs
[0136] 10 Data processing system 12 Data processing device 14 Smart device 290 Specific processing unit< / url:>
Claims
1. A data processing device comprising a processor, wherein the processor acquires electronic content in which content including text and illustrations is digitized, inputs a prompt including an image of a specific section among those partitioned in a predetermined partition unit in the content and an instruction to estimate the emotion of a character shown in the image of the specific section into a generation model that generates information according to input data, when an image of the specific section is displayed on a display unit capable of displaying the electronic content, causes a line of the character in the specific section to be output from an output unit in artificial voice generated based on the emotion of the character estimated by the generation model. A data processing device.
2. The processor when onomatopoeia is included in the image of the specific section, inputs a prompt including the image of the specific section and an instruction to interpret the onomatopoeia into the generation model, when the image of the specific section is displayed on the display unit, causes a sound effect generated based on the interpretation result of the onomatopoeia output by the generation model to be output from the output unit. The data processing device according to claim 1.
3. The processor when receiving a predetermined operation by a user while the image of the specific section is being displayed on the display unit, causes the estimated content of the emotion of the character generated by the generation model to be output from the output unit in a predetermined artificial voice. The data processing device according to claim 1.
4. The processor when the output of the sound corresponding to the image of the specific section from the output unit is completed, performs at least one of outputting a specific sound and generating a specific vibration using at least one of the sound output function of the output unit and the vibration function of a vibration unit. The data processing device according to claim 1.
5. The processor when the content is in video form, extracts a specific virtual actor corresponding to a specific actor in charge of the voice of the character from an actor database storing a plurality of virtual actors capable of outputting artificial voices with a predetermined voice quality, when the image of the specific section is displayed on the display unit, causes a line of the character in the specific section to be output from the output unit in artificial voice by the specific virtual actor. The data processing device according to claim 1.
6. The processor when there are a plurality of the specific actors, accepts a user's selection of one actor from among the plurality of the specific actors. Extract a first virtual actor corresponding to the one actor who has received the selection from the actor database. When an image of the specific section is displayed on the display unit, cause the output unit to output, by artificial voice of the first virtual actor, the lines of the character in the specific section. The data processing apparatus according to claim 5.
7. The processor When the content is not visualized, input the electronic content into the generation model and obtain the characteristics of the character interpreted thereby. Input a prompt including the obtained characteristics of the character and an instruction to seek an actor having a voice quality suitable for the characteristics into the generation model. Extract a second virtual actor corresponding to the actor output by the generation model from an actor database in which a plurality of virtual actors capable of outputting artificial voices of a predetermined voice quality are stored. When an image of the specific section is displayed on the display unit, cause the output unit to output, by artificial voice of the second virtual actor, the lines of the character in the specific section. The data processing apparatus according to claim 1.
8. Obtain electronic content in which content including text and illustrations is digitized. Input a prompt including an image of a specific section among those sectioned in units of a predetermined section in the content and an instruction to estimate the emotion of the character shown in the image of the specific section into a generation model that generates information according to input data. When an image of the specific section is displayed on a display unit capable of displaying the electronic content, cause the output unit to output, by artificial voice generated based on the emotion of the character estimated by the generation model, the lines of the character in the specific section. A data processing method for a computer to execute processing.
9. Obtain electronic content in which content including text and illustrations is digitized. Input a prompt including an image of a specific section among those sectioned in units of a predetermined section in the content and an instruction to estimate the emotion of the character shown in the image of the specific section into a generation model that generates information according to input data. When an image of the specific section is displayed on a display unit capable of displaying the electronic content, cause the output unit to output, by artificial voice generated based on the emotion of the character estimated by the generation model, the lines of the character in the specific section. A data processing program for causing a computer to execute processing.
Citation Information
Patent Citations
Selection-inference neural network systems
WO2023218040A1
Persona chatbot control method and system
JP2022180282A