Data processing device, data processing method, and data processing program

The data processing device improves user immersion in electronic content by using a generative model to estimate character emotions and interpret onomatopoeia, outputting emotionally nuanced speech and sound effects, addressing the limitations of existing generative models.

JP2025174291AActive Publication Date: 2025-11-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024080476
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-11-28
Estimated Expiration
2044-05-16

AI Technical Summary

Technical Problem

Existing generative models, such as language models, lack the capability to enhance the quality of artificial speech output by incorporating emotional intonation and sound effects, leading to a diminished user experience in interactive content.

Method used

A data processing device that inputs images of content sections and prompts to a generative model to estimate character emotions and interpret onomatopoeia, outputting the character's lines in an artificial voice and sound effects based on the model's estimation, and optionally using virtual actors to match voice qualities.

Benefits of technology

Enhances user immersion in electronic content by providing emotionally nuanced speech and sound effects, reducing incongruity and improving understanding, even for visually impaired users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025174291000001_ABST
    Figure 2025174291000001_ABST
Patent Text Reader

Abstract

To provide a data processing device, method, and program for enhancing quality of artificial speech output from a generative model based on electronic contents.SOLUTION: A data processing device includes a processor. The processor acquires electronic contents in which contents containing text and illustrations are digitized, inputs, to a generation model that generates information corresponding to input data, an image (frame image 70A) of a specific section among sections divided by a predetermined one section unit in the contents and a prompt 70 that includes an instruction sentence 70B indicating an instruction to estimate feeling of a character shown on the image of the specific section, and causes an output unit to output lines of the character in the specific section with an artificial voice generated based on the feeling of the character estimated by the generation model, when the image of the specific section is displayed on a display unit capable of displaying the electronic contents.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a data processing device, a data processing method, and a data processing program. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the prior art, there is still room for improvement in the quality of artificial speech output from a generative model such as the above-mentioned language model. [Means for solving the problem]

[0005] A first aspect of the data processing device includes a processor that acquires electronic content in which content including text and illustrations has been digitized, and inputs an image of a specific section of the content divided into predetermined units of sections, and a prompt including instructions to estimate the emotion of a character shown in the image of the specific section, into a generative model that generates information according to input data, and when the image of the specific section is displayed on a display unit capable of displaying the electronic content, the lines of the character in the specific section are output from an output unit in an artificial voice generated based on the emotion of the character estimated by the generative model.

[0006] In a first aspect of the data processing device, a processor acquires electronic content. A prompt including an image of a specific section in the electronic content and an instruction to estimate the emotion of a character depicted in the image of the specific section is input to a generative model. When the image of the specific section is displayed on a display unit, an output unit outputs lines of the character in the specific section in an artificial voice generated based on the emotion of the character estimated by the generative model. This allows the data processing device to enhance a user's sense of immersion in the electronic content compared to a configuration in which the output unit outputs an artificial voice with no intonation.

[0007] In the data processing device of the second aspect, in the first aspect, when the image of the specific section includes an onomatopoeia, the processor inputs the image of the specific section and a prompt including instructions to interpret the onomatopoeia into the generative model, and when the image of the specific section is displayed on the display unit, outputs from the output unit a sound effect generated based on the interpretation result of the onomatopoeia output by the generative model.

[0008] In the data processing device of the second aspect, when an image of a specific segment includes an onomatopoeia, the image of the specific segment and a prompt including an instruction to interpret the onomatopoeia are input to the generative model. Then, when the image of the specific segment is displayed on the display unit, a sound effect generated based on the interpretation result of the onomatopoeia output by the generative model is output from the output unit. This allows the data processing device to enhance the user's sense of immersion in the electronic content compared to a configuration in which sound effects corresponding to the onomatopoeia are not output from the output unit.

[0009] In the data processing device of the third aspect, in the first or second aspect, when the processor receives a predetermined operation from the user while an image of the specific area is displayed on the display unit, the processor outputs the estimated emotion of the character generated by the generative model from the output unit in a predetermined artificial voice.

[0010] In the data processing device of the third aspect, when a predetermined operation is received from the user while an image of a specific section is displayed on the display unit, the output unit outputs an estimated emotion of the character generated by the generative model in a predetermined voice. This allows the data processing device to improve the user's understanding of the content of the electronic content compared to a configuration in which only the character's lines are output as voice.

[0011] A fourth aspect of the data processing device is any one of the first to third aspects, in which, when the output of sound corresponding to the image of the specific section from the output unit has ended, the processor performs at least one of outputting a specific sound and generating a specific vibration using at least one of the sound output function of the output unit and the vibration function of the vibration unit.

[0012] In the data processing device of the fourth aspect, when the output of the sound corresponding to the image of the specific section from the output unit is completed, at least one of the sound output function and the vibration function is used to output the specific sound and / or generate the specific vibration. In this way, even if the user is visually impaired, the data processing device allows the user to know that it is time to switch the section displayed on the display unit.

[0013] A fifth aspect of the data processing device is any one of the first to fourth aspects, wherein when the content is visualized, the processor extracts a specific virtual actor corresponding to a specific actor who voiced the character from an actor database that stores a plurality of virtual actors capable of outputting artificial voices with a predetermined voice quality, and when an image of the specific section is displayed on the display unit, the lines of the character in the specific section are output from the output unit in the artificial voice of the specific virtual actor.

[0014] In the data processing device of the fifth aspect, when content is visualized, a specific virtual actor corresponding to a specific actor who voiced a character is extracted from the actor database. Then, when an image of a specific section is displayed on the display unit, the output unit outputs the lines of the character in the specific section in an artificial voice uttered by the specific virtual actor. This makes it possible to reduce the sense of incongruity felt by a user who is familiar with the voice of the character when visualized, due to the artificial voice of the character.

[0015] A sixth aspect of the data processing device is the fifth aspect, wherein, when there are multiple specific actors, the processor accepts a user's selection of one of the multiple specific actors, extracts a first virtual actor corresponding to the one actor whose selection was accepted from the actor database, and when an image of the specific section is displayed on the display unit, outputs the lines of the character in the specific section from the output unit in an artificial voice by the first virtual actor.

[0016] In a data processing device of a sixth aspect, when there are multiple specific actors, the user selects one of the multiple specific actors. A first virtual actor corresponding to the one actor selected by the user is extracted from the actor database. Then, when an image of the specific section is displayed on the display unit, the lines of the character in the specific section are output from the output unit using an artificial voice uttered by the first virtual actor. Thus, with this data processing device, when there are multiple specific actors, it is possible to set an artificial voice for the character that suits the user's preferences.

[0017] A seventh aspect of the data processing device is any one of the first to sixth aspects, in which, when the content is not visualized, the processor inputs the electronic content into the generative model to obtain interpreted characteristics of the character, inputs into the generative model a prompt including the obtained characteristics of the character and an instruction to find an actor with a voice quality suitable for the characteristics, extracts a second virtual actor corresponding to the actor output by the generative model from an actor database storing a plurality of virtual actors capable of outputting artificial voices with a predetermined voice quality, and, when an image of the specific section is displayed on the display unit, outputs the lines of the character in the specific section from the output unit using the artificial voice of the second virtual actor.

[0018] In a data processing device of a seventh aspect, when content has not been visualized, electronic content is input into a generative model to obtain interpreted character characteristics. A prompt including the character characteristics and an instruction to find an actor with a voice quality suitable for the characteristics is input into the generative model. A virtual actor corresponding to the actor output by the generative model is extracted from the actor database. Then, when an image of a specific segment is displayed on the display unit, the output unit outputs the lines of the character in the specific segment in an artificial voice uttered by the virtual actor. As a result, with this data processing device, even if the content has not been visualized, it is possible to set an artificial voice for the character that can reproduce a voice quality suitable for the character's characteristics.

[0019] An eighth aspect of the data processing method involves a computer executing a process in which electronic content is obtained in which content including text and illustrations has been digitized, and inputting an image of a specific section of the content divided into predetermined sections, and a prompt including instructions to estimate the emotion of a character shown in the image of the specific section, into a generative model that generates information according to input data, and when the image of the specific section is displayed on a display unit capable of displaying the electronic content, the lines of the character in the specific section are output from an output unit in an artificial voice generated based on the emotion of the character estimated by the generative model.

[0020] In an eighth aspect of the data processing method, electronic content is acquired. A prompt including an image of a specific section in the electronic content and an instruction to estimate the emotion of a character shown in the image of the specific section is input to a generative model. When the image of the specific section is displayed on the display unit, the output unit outputs lines of the character in the specific section in an artificial voice generated based on the emotion of the character estimated by the generative model. This data processing method can enhance the user's sense of immersion in the electronic content compared to a configuration in which the output unit outputs an artificial voice with no intonation.

[0021] A ninth aspect of the data processing program causes a computer to execute the following process: acquire electronic content in which content including text and illustrations has been digitized; input an image of a specific section of the content, which is divided into predetermined sections, and a prompt including instructions to estimate the emotion of a character shown in the image of the specific section, into a generative model that generates information according to the input data; and, when the image of the specific section is displayed on a display unit capable of displaying the electronic content, output the lines of the character in the specific section from an output unit in an artificial voice generated based on the emotion of the character estimated by the generative model.

[0022] In a ninth aspect of the data processing program, electronic content is acquired. A prompt including an image of a specific section in the electronic content and an instruction to estimate the emotion of a character shown in the image of the specific section is input to a generative model. When the image of the specific section is displayed on the display unit, the output unit outputs lines of the character in the specific section in an artificial voice generated based on the emotion of the character estimated by the generative model. This data processing program can enhance the user's sense of immersion in the electronic content compared to a configuration in which the output unit outputs an artificial voice with no intonation. [Brief explanation of the drawings]

[0023] [Figure 1] FIG. 1 is a conceptual diagram illustrating an example of a configuration of a data processing system. [Figure 2] FIG. 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device. [Figure 3] FIG. 10 is a first explanatory diagram showing an example of a specific process performed by a data processing device. [Figure 4] FIG. 2 is a second explanatory diagram showing an example of the specific processing performed by the data processing device. [Figure 5] 10 is a diagram illustrating an example of an operational flow of a first identification process performed by the data processing device. [Figure 6] This is a subroutine of the frame information generation process. [Figure 7] This is a subroutine of emotion estimation processing. [Figure 8] This is a subroutine for sound effect generation processing. [Figure 9] 10 is a subroutine of the artificial voice determination process. [Figure 10] 10 shows an example of an operational flow of a second identification process by the data processing device. [Figure 11] FIG. 10 is a first explanatory diagram showing a display example of a display. [Figure 12] FIG. 2 is a second explanatory diagram showing a display example of the display. [Figure 13] FIG. 10 is a third explanatory diagram showing a display example of the display. DETAILED DESCRIPTION OF THE INVENTION

[0024] Hereinafter, an example of an embodiment of a data processing device, a data processing method, and a program according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0025] First, the terms used in the following description will be explained.

[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), or an APU (Accelerated Processing Unit).

[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the embodiment.

[0032] As shown in FIG. 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server. An example of the smart device 14 is a smartphone. In this embodiment, the data processing device 12 is an example of a "data processing device" according to the technology of the present disclosure.

[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28 is an example of a "processor" according to the technology of the present disclosure. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The display 40A is an example of a "display unit" according to the technology of the present disclosure, and the speaker 40B is an example of an "output unit" according to the technology of the present disclosure.

[0037] The camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, a shutter, and the like, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0040] As shown in FIG. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "data processing program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30. The "data processing program" according to the technology of the present disclosure can also be applied as a program product.

[0041] The storage 32 stores a data generation model 58. The data generation model 58 is used by the specific processing unit 290.

[0042] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">), EmotiVoice (Internet search<URL: https: / / weel.co.jp / media / tech / emotivoice / > ), and Audiobox (Internet search<URL: https: / / audiobox.metademolab.com / > ) and other generative AIs. The data generation model 58 is configured by appropriately combining various known generative AIs such as those described above. The data generation model 58 is obtained by causing a neural network to perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The data generation model 58 is an example of a "generative model" according to the technology of the present disclosure.

[0043] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0044] Next, an example of the specific processing performed by the data processing device 12 will be described. FIG. 3 is a first explanatory diagram showing an example of the identification process performed by the data processing device 12. FIG. 3 shows an example in which, in the identification process, a data generation model 58 estimates the emotion of a character shown in an image of a specific frame (hereinafter simply referred to as a "frame") of an electronic comic. The data generation model 58 is, for example, ChatGPT. An electronic comic is a paper comic (hereinafter simply referred to as a "comic") that includes text and illustrations and has been digitized. In the comic, each page is divided into predetermined frame units. An electronic comic is an example of "electronic content" according to the technology of the present disclosure, a comic is an example of "content" according to the technology of the present disclosure, and a frame unit is an example of a "single section unit" according to the technology of the present disclosure.

[0045] 3 shows a prompt 70 to be input to the data generation model 58. The prompt 70 includes a frame image 70A showing an image of a specific frame in the comic A as the comic, and an instruction statement 70B showing an instruction to estimate the emotion of the character shown in the frame image 70A. The frame image 70A is an example of an "image of a specific section" according to the technology of the present disclosure.

[0046] Frame image 70A shows a character saying, "I wonder if it will be sunny tomorrow." In addition, instruction sentence 70B is, for example, text that reads, "How do you think the character in this frame is feeling?"

[0047] 3 also shows an output result 71 when a prompt 70 is input to the data generation model 58. An example of the output result 71 is the text "The image shows a teary-eyed character standing still. From this, it can be inferred that the character is feeling anxious or sad."

[0048] Fig. 4 is a second explanatory diagram showing an example of the identification process performed by the data processing device 12. Fig. 4 shows an example in which the data generation model 58 interprets onomatopoeia shown in an image of a specific frame of a digital comic in the identification process. The data generation model 58 is, for example, ChatGPT.

[0049] 4 shows a prompt 72 to be input to the data generation model 58. The prompt 72 includes a frame image 72A showing an image of a specific frame in comic A, and instructional text 72B showing an instruction for interpreting the onomatopoeia shown in the frame image 72A. The frame image 72A is an example of an "image of a specific section" according to the technology of the present disclosure.

[0050] The frame image 72A shows the onomatopoeia "Dogyuun." The instruction sentence 72B is, for example, the text "What sound does the onomatopoeia in this frame make?"

[0051] FIG. 4 also shows an output result 73 when a prompt 72 is input to the data generation model 58. One example of the output result 73 is the text, "The onomatopoeia 'Dogyun' is an onomatopoeia that indicates a large impact or an object moving at high speed."

[0052] Next, the operation of the data processing system 10 will be described. An example of the flow of the specification process will be described with reference to FIGS. 5 to 10. The specification process includes a first specification process shown in FIGS. 5 to 9 and a second specification process shown in FIG. 10. As an example, the first specification process is performed when a user operates the smart device 14 to execute a predetermined application installed on the smart device 14, and a selection screen for selecting an e-comic is displayed on the display 40A. The second specification process is performed when a user operates the smart device 14 to execute the predetermined application, and a viewing screen for viewing the e-comic is displayed on the display 40A. The flow of the specification process shown in FIGS. 5 to 10 is an example of a "data processing method" according to the technology of the present disclosure.

[0053] In step S10 shown in Fig. 5, the processor 28 retrieves a specific e-comic from the database 24 based on data indicating a user input to the smart device 14, for example, a user input by voice. The database 24 stores various e-comics that are digitized versions of various manga. Hereinafter, the specific e-comic retrieved from the database 24 will be described as "e-comic A," which is a digitized version of manga A. Then, the processor 28 proceeds to step S11.

[0054] In step S11, processor 28 lists the characters appearing in digital comic A. Here, processor 28 inputs digital comic A, instructions for interpreting the story and style of digital comic A, and a prompt including instructions for listing the characteristics of each character appearing in digital comic A to data generation model 58, and obtains the output result. The data generation model 58 is, for example, ChatGPT. The prompt is, for example, text and PDF data such as "Please interpret the story and style of the following digital comic A. Also, please list the characteristics of each character appearing in digital comic A. Digital Comic A.pdf." As a result, in step S11, based on the output result from data generation model 58, the characteristics of each character are listed, such as "Character A: A boy about 10 years old with a quiet personality; Character B: A boy about 10 years old with a lively personality..." Then, processor 28 proceeds to step S12.

[0055] In step S12, processor 28 determines background music (BGM) to be output from speaker 40B while electronic comic A is being displayed on display 40A. Database 24 stores various types of BGM corresponding to the styles of various electronic comics. Processor 28 obtains, from database 24, BGM corresponding to the style of electronic comic A indicated in the output result by data generation model 58 in step S11. Processor 28 transmits sound data indicating the obtained BGM to processor 46. Processor 28 then proceeds to step S13.

[0056] In step S13, processor 28 performs a frame information generation process to generate frame information to be included in each frame of electronic comic A. The frame information includes information indicating the frame image as well as various other information, which will be described later. A subroutine for the frame information generation process will be described later. Processor 28 then proceeds to step S14.

[0057] In step S14, processor 28 performs emotion estimation processing to estimate the emotion of a character appearing in each frame of digital comic A. A subroutine of the emotion estimation processing will be described later. Processor 28 then proceeds to step S15.

[0058] In step S15, processor 28 performs a sound effect generation process to generate a sound effect corresponding to the onomatopoeia shown in a specific frame of electronic comic A. A subroutine of the sound effect generation process will be described later. Processor 28 then proceeds to step S16.

[0059] In step S16, processor 28 performs an artificial voice determination process to determine an artificial voice to be used to utter lines for characters appearing in digital comic A. A subroutine of the artificial voice determination process will be described later. Then, processor 28 ends the process.

[0060] FIG. 6 shows a subroutine for the frame information generation process. 6, processor 28 divides electronic comic A into frames. Processor 28 then proceeds to step S21. As an example, the order of frames in electronic comic A is predetermined, and processor 28 performs the processes from step S21 onward in the predetermined frame order.

[0061] In step S21, processor 28 identifies the characters appearing in the frame. For example, processor 28 identifies the characters appearing in the frame by referring to the image of the frame and the characteristics of each character listed in step S11. Processor 28 then proceeds to step S22.

[0062] In step S22, processor 28 performs character recognition processing on the frame image to identify the lines of the characters appearing in the frame. If multiple characters appear in the frame image and multiple lines are present, processor 28 identifies the speaker of each line based on the image information, character recognition information, etc. Then, processor 28 proceeds to step S23.

[0063] In step S23, processor 28 associates the lines of the character identified in step S21 and the character identified in step S22 with the frame information of the corresponding frame and stores them in database 24. As a result, information is stored in database 24 that, for example, character A appears in the first frame of electronic comic A and that character A's line is "I wonder if it will be sunny tomorrow." Processor 28 then proceeds to step S24.

[0064] In step S24, processor 28 determines whether the frame information stored in database 24 in step S23 corresponds to the last frame of digital comic A. If processor 28 determines that the frame information corresponds to the last frame of digital comic A (step S24: YES), it returns to the process that called the process. On the other hand, if processor 28 determines that the frame information does not correspond to the last frame of digital comic A (step S24: NO), it proceeds to step S25. As an example, a specific label is assigned to the last frame of digital comic A, and processor 28 determines whether the frame information corresponds to the last frame based on the presence or absence of the specific label.

[0065] In step S25, processor 28 advances the processing target to the next frame. Then, processor 28 returns to step S21. In this manner, processor 28 repeatedly executes the subroutine shown in FIG. 6 from the first frame to the last frame of electronic comic A.

[0066] FIG. 7 shows a subroutine of the emotion estimation process. 7, processor 28 obtains frame information corresponding to the first frame of digital comic A from database 24. As an example, a predetermined label is assigned to the first frame of digital comic A, and processor 28 obtains frame information corresponding to the frame to which the predetermined label is assigned from database 24. Processor 28 then proceeds to step S31.

[0067] In step S31, processor 28 generates a prompt to be input to data generation model 58. Data generation model 58 is, for example, ChatGPT. The prompt includes an image of a frame of digital comic A to be processed and an instruction to estimate the emotion of the character depicted in the image. The frame to be processed is the first frame the first time the subroutine shown in FIG. 7 is executed, and from the second time onwards, it is a frame corresponding to the frame information acquired in a step within the processing. The "frame to be processed" that appears hereinafter has the same meaning. For example, the prompt includes the image shown in frame image 70A in FIG. 3 and the text shown in instruction statement 70B. Processor 28 then proceeds to step S32.

[0068] In step S32, the processor 28 inputs the prompt generated in step S31 to the data generation model 58 and obtains an output result from the data generation model 58. Then, the processor 28 proceeds to step S33.

[0069] In step S33, processor 28 associates the output result output from data generation model 58 in step S32 with the frame information of the frame to be processed and stores the result in database 24. For example, the output result is text such as that shown in output result 71 in FIG. 3. Processor 28 then proceeds to step S34.

[0070] In step S34, processor 28 determines whether the frame information stored in database 24 in step S33 corresponds to the last frame of electronic comic A. If processor 28 determines that the frame information corresponds to the last frame of electronic comic A (step S34: YES), it returns to the calling process. On the other hand, if processor 28 determines that the frame information does not correspond to the last frame of electronic comic A (step S34: NO), it proceeds to step S35.

[0071] In step S35, processor 28 advances the processing target to the next frame and obtains frame information corresponding to the next frame from database 24. Processor 28 then returns to step S31. In this manner, processor 28 repeatedly executes the subroutine shown in FIG. 7 from the first frame to the last frame of electronic comic A.

[0072] FIG. 8 shows a subroutine of the sound effect generation process. 8, processor 28 obtains frame information corresponding to the first frame of electronic comic A from database 24. Processor 28 then proceeds to step S41.

[0073] In step S41, processor 28 determines whether or not an onomatopoeia is included in the frame to be processed. If processor 28 determines that an onomatopoeia is included (step S41: YES), the process proceeds to step S42. On the other hand, if processor 28 determines that an onomatopoeia is not included (step S41: NO), the process proceeds to step S47. Database 24 stores various onomatopoeia used in various manga. Processor 28 performs character recognition processing on the image of the frame to be processed, and determines whether or not an onomatopoeia is included based on whether or not the result of the character recognition processing matches an onomatopoeia stored in database 24.

[0074] In step S42, processor 28 generates a prompt to be input to data generation model 58. Data generation model 58 is, for example, ChatGPT. The prompt includes an image of the frame to be processed and an instruction to interpret the onomatopoeia shown in the image. For example, the prompt includes the image shown in frame image 72A in FIG. 4 and the text shown in instruction statement 72B. Processor 28 then proceeds to step S43.

[0075] In step S43, the processor 28 inputs the prompt generated in step S42 into the data generation model 58 and obtains an output result from the data generation model 58. The processor 28 then proceeds to step S44.

[0076] In step S44, processor 28 generates a prompt to be input to data generation model 58. Data generation model 58 is, for example, an Audiobox. The prompt includes an interpretation result of the onomatopoeia, which is the output result of data generation model 58 in step S43, and an instruction to generate a sound effect according to the interpretation result of the onomatopoeia. For example, the interpretation result of the onomatopoeia is text such as that shown in output result 73 in FIG. 4. As a result, the prompt becomes text such as, "The onomatopoeia 'dugyuun' is an onomatopoeia that indicates a large impact or an object moving at high speed. Please generate a sound effect that is appropriate for this onomatopoeia." Processor 28 then proceeds to step S45.

[0077] In step S45, the processor 28 inputs the prompt generated in step S44 into the data generation model 58 and obtains an output result from the data generation model 58. The processor 28 then proceeds to step S46.

[0078] In step S46, processor 28 associates the output result output from data generation model 58 in step S45 with the frame information of the frame to be processed and stores the result in database 24. For example, the output result may be sound data indicating a sound effect. Processor 28 then proceeds to step S47.

[0079] In step S47, processor 28 determines whether the frame information stored in database 24 in step S46 corresponds to the last frame of digital comic A. If processor 28 determines that the frame information corresponds to the last frame of digital comic A (step S47: YES), it returns to the calling process. On the other hand, if processor 28 determines that the frame information does not correspond to the last frame of digital comic A (step S47: NO), it proceeds to step S48.

[0080] In step S48, processor 28 advances the processing target to the next frame and obtains frame information corresponding to the next frame from database 24. Processor 28 then returns to step S41. In this manner, processor 28 repeatedly executes the subroutine shown in FIG. 7 from the first frame to the last frame of electronic comic A.

[0081] FIG. 9 shows a subroutine of the artificial voice determination process. 9, processor 28 selects any one character from among the characters appearing in digital comic A. Processor 28 then proceeds to step S51.

[0082] In step S51, processor 28 determines whether manga A has been animated. If processor 28 determines that manga A has been animated (step S51: YES), the process proceeds to step S52. On the other hand, if processor 28 determines that manga A has not been animated (step S51: NO), the process proceeds to step S54. As an example, processor 28 determines whether manga A has been animated based on the output result of data generation model 58. The data generation model 58 is, for example, ChatGPT. In this case, processor 28 inputs a prompt to data generation model 58 asking whether manga A has been animated, such as "Has manga A been animated?", and obtains the output result from data generation model 58. Animation is an example of "visualization" according to the technology of the present disclosure.

[0083] In step S52, processor 28 identifies the voice actor who voiced the target character in the animation. The target character is the character selected in step S50 the first time the subroutine shown in FIG. 9 is executed, and the character selected in step S59 the second time and thereafter. The "target character" that appears hereinafter has the same meaning. As an example, processor 28 identifies the voice actor based on the output result of data generation model 58. The data generation model 58 is, for example, ChatGPT. In this case, processor 28 inputs a prompt to data generation model 58, such as "Who is the voice actor who voiced character A in the animation of manga A?", to ask about the voice actor who voiced the target character in the animation of manga A, and obtains the output result from data generation model 58. Processor 28 then proceeds to step S53.

[0084] In step S53, the processor 28 determines a virtual voice actor who will deliver the lines of the character being processed. The database 24 stores a voice actor database that stores a plurality of virtual voice actors capable of outputting artificial voices with predetermined voice qualities. A virtual voice actor is a virtual voice actor that has been trained based on the voice of a real voice actor so that it can produce an artificial voice similar to that of the real voice actor. Examples of predetermined voice qualities include a clear, strong voice, a childish, high-pitched voice, a lively and bright voice, a gentle and cute voice, a low, cool voice, a refreshing young man's voice, a boy's voice immediately after puberty, a deep, low-pitched voice, a soft and warm voice, and an elegant adult voice. The voice actor database stores, linked to each virtual voice actor, the voice actor corresponding to the virtual voice actor (in other words, the voice actor on whom the artificial voice of the virtual voice actor was based), the characters that the voice actor has voiced, and the voice qualities corresponding to the virtual voice actor. As a result, the voice actor database stores information such as, for example, "The voice actor corresponding to virtual voice actor A is voice actor A, the characters that voice actor A has voiced are character A, character C, and character E, etc., and the voice quality corresponding to virtual voice actor A is a clear, strong voice."

[0085] Here, processor 28 extracts from the voice actor database a specific virtual voice actor corresponding to the voice actor identified in step S52. As a result, if processor 28 identified voice actor A in step S52, processor 28 determines virtual voice actor A corresponding to voice actor A as the virtual voice actor who will speak the lines of the character being processed. Processor 28 then proceeds to step S57. The voice actor database is an example of an "actor database" according to the technology of the present disclosure, the voice actor is an example of an "actor" according to the technology of the present disclosure, the virtual voice actor is an example of a "virtual actor" according to the technology of the present disclosure, and virtual voice actor A is an example of a "specific virtual actor" according to the technology of the present disclosure.

[0086] In step S54, processor 28 generates a prompt to be input to data generation model 58. Data generation model 58 is, for example, ChatGPT. The prompt includes the characteristics of the character to be processed and an instruction to ask for a voice actor with a voice quality that suits the characteristics. The characteristics of the character to be processed use the information listed in step S11 shown in FIG. 5. As a result, the prompt becomes text such as, for example, "Character B is a boy about 10 years old with a lively personality. Which voice actor has a voice quality that suits the characteristics of this character B?" Then, processor 28 proceeds to step S55.

[0087] In step S55, the processor 28 inputs the prompt generated in step S54 into the data generation model 58 and obtains an output result from the data generation model 58. The processor 28 then proceeds to step S56.

[0088] In step S56, processor 28 determines a virtual voice actor who will speak the lines of the character being processed. Here, processor 28 extracts from the voice actor database a virtual voice actor corresponding to the voice actor indicated in the output result of data generation model 58 in step S55. As a result, if the output result of data generation model 58 is voice actor B, processor 28 determines virtual voice actor B corresponding to voice actor B as the virtual voice actor who will speak the lines of the character being processed. Processor 28 then proceeds to step S57. Virtual voice actor B is an example of a "second virtual actor" according to the technology of the present disclosure.

[0089] In step S57, processor 28 associates the character to be processed with the virtual voice actor who will utter the lines of that character, and stores them in database 24. Processor 28 then proceeds to step S58.

[0090] In step S58, processor 28 determines whether or not the association of all characters with virtual voice actors has been completed. If processor 28 determines that the association of all characters with virtual voice actors has been completed (step S58: YES), the process returns to the calling process. On the other hand, if processor 28 determines that the association of all characters with virtual voice actors has not been completed (step S58: NO), the process proceeds to step S59.

[0091] In step S59, processor 28 selects the next character to be processed. Then, processor 28 returns to step S51. In this manner, processor 28 repeatedly executes the subroutine shown in FIG. 9 until all characters appearing in digital comic A have been linked to virtual voice actors.

[0092] FIG. 10 is a flowchart showing the flow of the second identification process. 10, processor 28 obtains frame information corresponding to the frame designated by the user from database 24. Processor 28 then proceeds to step S61.

[0093] In step S61, processor 28 transmits to smart device 14 piece information corresponding to the piece specified by the user in step S60. As an example, processor 28 transmits at least the piece information, including an image of the piece, character recognition results for the dialogue of the character in the piece, and an estimation result for the emotion of the character appearing in the piece. If the piece includes an onomatopoeia, processor 28 additionally transmits sound data indicating a sound effect corresponding to the onomatopoeia. As a result, an image of the piece is displayed on display 40A of smart device 14. Furthermore, processor 46 of smart device 14 outputs background music indicated in the acquired sound data from speaker 40B based on the display of the image of the piece. Then, processor 28 proceeds to step S62.

[0094] In step S62, the processor 28 instructs the smart device 14 to select a virtual voice actor to deliver the lines of the character appearing in the current frame based on the frame information acquired in step S60. The processor 28 determines the virtual voice actor to deliver the lines of the character appearing in the current frame based on the frame information and the association between each character and each virtual voice actor stored in the database 24, and notifies the processor 46 of the determined virtual voice actor. The storage 50 of the smart device 14 stores text-to-speech software capable of reading text aloud in the artificial voice of each virtual voice actor registered in the voice actor database. The processor 46 of the smart device 14, in accordance with instructions from the processor 28, selects a virtual voice actor to deliver the lines of the current frame from the text-to-speech software. The processor 46 then uses the text-to-speech software to generate sound data in which the selected virtual voice actor reads the text aloud in an artificial voice based on the estimated emotion of the character included in the frame information transmitted from the processor 28, and instructs the speaker 40B to output the sound data. As a result, the speaker 40B outputs the lines of the character in the frame in the artificial voice of the set virtual voice actor. Furthermore, since the virtual voice actor reads the text based on the estimation result of the character's emotion included in the frame information transmitted from the processor 28, the speaker 40B outputs the character's lines in an artificial voice that reflects the character's emotion estimated by the data generation model 58. Furthermore, if the frame includes an onomatopoeia, the speaker 40B outputs a sound effect corresponding to the onomatopoeia indicated in the acquired sound data. Although detailed description is omitted here, when the output of the artificial voice indicating the lines of the character in the frame or the sound effect corresponding to the onomatopoeia has ended, the processor 46 of the smart device 14 outputs a specific sound from the speaker 40B similar to that in step S66, which will be described later. Then, the processor 28 proceeds to step S63.

[0095] In step S63, processor 28 determines whether or not there has been a change in the frame of electronic comic A displayed on display 40A. If processor 28 determines that there has been a change in the frame (step S63: YES), the process proceeds to step S67. On the other hand, if processor 28 determines that there has not been a change in the frame (step S63: NO), the process proceeds to step S64. As an example, processor 28 determines that there has been a change in the frame when it acquires data indicating a user input by a flick operation on display 40A. The data indicating the user input on display 40A is transmitted from smart device 14 to data processing device 12 as appropriate.

[0096] In step S64, processor 28 determines whether a predetermined operation has been performed on display 40A. If processor 28 determines that a predetermined operation has been performed (step S64: YES), the process proceeds to step S65. On the other hand, if processor 28 determines that a predetermined operation has not been performed (step S64: NO), the process proceeds to step S68. As an example, processor 28 determines that a predetermined operation has been performed when processor 28 acquires data indicating a user input by a tap operation on display 40A.

[0097] In step S65, the processor 28 instructs the smart device 14 to output a commentary for the piece being processed. The commentary for the piece includes an estimated emotion of the character included in the piece information, an estimated state of the character based on an analysis result of the image of the piece, and the like. As a result, the commentary for the piece is output from the speaker 40B in a predetermined artificial voice. The commentary for the piece includes, for example, text content such as that shown in the output result 71 in FIG. 3 as an estimated emotion of the character. The predetermined artificial voice may be an artificial voice by any virtual voice actor in the text-to-speech software, or may be an artificial voice imitating the voice of a person (e.g., mother or father) designated in advance by the user. Note that if the predetermined artificial voice is an artificial voice imitating the voice of a person designated in advance by the user, the smart device 14 is configured in advance to enable output of the artificial voice from the speaker 40B. The processor 28 then proceeds to step S66.

[0098] In step S66, the processor 28 instructs the smart device 14 to output a specific sound. The specific sound is a sound corresponding to the image of the frame being processed, such as an artificial voice providing an explanation for the frame, an artificial voice providing lines from the character in the frame, or a sound for notifying the user that output of a sound effect corresponding to the onomatopoeia for the frame has ended. The type of the specific sound is not particularly limited. As a result, the specific sound is output from the speaker 40B. The processor 28 then returns to step S63. Note that data indicating whether or not sound is being output from the speaker 40B is appropriately transmitted from the smart device 14 to the data processing device 12.

[0099] In step S67, processor 28 advances the frame to be processed to the next frame, and then processor 28 returns to step S60.

[0100] In step S68, processor 28 determines whether or not a termination condition for a predetermined application is satisfied. If processor 28 determines that the termination condition is satisfied (step S68: YES), it terminates the process. On the other hand, if processor 28 determines that the termination condition is not satisfied (step S68: NO), it returns to step S63. As an example, processor 28 determines that the termination condition is satisfied when it acquires data indicating a user input corresponding to a termination operation for terminating the predetermined application.

[0101] Next, an example of data output from the output device 40 of the smart device 14 based on the execution of a specific process will be described.

[0102] Fig. 11 is a first explanatory diagram showing a display example of the display 40A. The display 40A shown in Fig. 11 displays a frame image 80 showing the image of the first frame of electronic comic A. The frame image 80 includes a character C1 and a line 80A spoken by the character C1. The character C1 is a human character. The content of the line 80A is "I wonder if it will be sunny tomorrow."

[0103] At this time, based on the frame image 80 being displayed on the display 40A, the content of the text shown in the dialogue 80A is output from the speaker 40B in the artificial voice of the virtual voice actor (e.g., virtual voice actor E) corresponding to the character C1. Furthermore, since the virtual voice actor E reads out the text based on the estimation result of the emotion of the character C1, the artificial voice that reflects the emotion of the character C1 is output from the speaker 40B.

[0104] Fig. 12 is a second explanatory diagram showing a display example of display 40A. Display 40A shown in Fig. 12 displays frame image 81 showing the second frame of digital comic A. Frame image 81 includes onomatopoeia 81A. Onomatopoeia 81A is an onomatopoeic word meaning "dogyuun."

[0105] At this time, based on the frame image 81 being displayed on the display 40A, a sound effect generated by the data generation model 58 in response to the onomatopoeia 81A is output from the speaker 40B. The data generation model 58 is, for example, an audio box.

[0106] Fig. 13 is a third explanatory diagram showing a display example of the display 40A. The display 40A shown in Fig. 13 displays a frame image 82 showing the third frame of electronic comic A. The frame image 82 includes a character C1, a character C2, an onomatopoeia 82A, a line 82B for character C2, and a line 82C for character C1. The character C2 is a human character. The onomatopoeia 82A is an onomatopoeic word meaning "ta-da." The content of the line 82B is "It's sure to clear up!" The content of the line 82C is "That's right, thank you!"

[0107] At this time, based on the frame image 82 being displayed on the display 40A, sounds corresponding to the frame images 82 are output from the speaker 40B in a predetermined order. In this embodiment, when there are multiple sounds corresponding to the frame images, the sounds are output from the speaker 40B in a predetermined output order. As an example, for the third frame image, the corresponding sounds are output in the order of onomatopoeia 82A, dialogue 82B, and dialogue 82C.

[0108] As a result, based on the display of the frame image 82 on the display 40A, first, a sound effect generated by the data generation model 58 corresponding to the onomatopoeia 82A is output from the speaker 40B. The data generation model 58 is, for example, an audio box. Next, the content of the text shown in the dialogue 82B is output from the speaker 40B in the artificial voice of the virtual voice actor (e.g., virtual voice actor B) corresponding to the character C2. Furthermore, since the virtual voice actor B reads the text based on the estimation result of the emotion of the character C2, the artificial voice reflecting the emotion of the character C2 is output from the speaker 40B. Finally, the content of the text shown in the dialogue 82C is output from the speaker 40B in the artificial voice of the virtual voice actor (e.g., virtual voice actor E) corresponding to the character C1. Furthermore, since the virtual voice actor E reads the text based on the estimation result of the emotion of the character C1, the artificial voice reflecting the emotion of the character C1 is output from the speaker 40B.

[0109] As described above, in the data processing device 12, the processor 28 acquires digital comic A, which is a digitalized version of manga A. The processor 28 also inputs to the data generation model 58 an image of a specific frame of manga A, which is divided into predetermined frame units, and a prompt including an instruction to estimate the emotion of a character depicted in the image of the specific frame. When the image of the specific frame is displayed on the display 40A, the processor 28 then outputs the lines of the character in the specific frame from the speaker 40B in an artificial voice generated based on the emotion of the character estimated by the data generation model 58. This allows the data processing device 12 to enhance the user's immersion in digital comic A compared to a configuration in which an artificial voice with no intonation is output from the speaker 40B. The specific frame is an example of a "specific section" according to the technology of the present disclosure.

[0110] Furthermore, in the data processing device 12, when an image of a specific frame contains an onomatopoeia, the processor 28 inputs the image of the specific frame and a prompt including an instruction to interpret the onomatopoeia to the data generation model 58. Then, when the image of the specific frame is displayed on the display 40A, the processor 28 causes the speaker 40B to output a sound effect generated based on the interpretation result of the onomatopoeia output by the data generation model 58. This allows the data processing device 12 to enhance the user's sense of immersion in the electronic comic A compared to a configuration in which sound effects corresponding to the onomatopoeia are not output from the speaker 40B.

[0111] Furthermore, in the data processing device 12, when a predetermined operation is received from the user while an image of a specific frame is being displayed on the display 40A, the processor 28 causes the speaker 40B to output a frame commentary, including estimated details of the character's emotions generated by the data generation model 58, in a predetermined artificial voice. This allows the data processing device 12 to improve the user's understanding of the content of the digital comic A compared to a configuration in which only the character's lines are output as voice.

[0112] Furthermore, in the data processing device 12, when the output of the sound corresponding to the image of a specific frame from the speaker 40B has finished, the processor 28 outputs a specific sound using the sound output function of the speaker 40B. The sound corresponding to the image of the specific frame is at least one of an artificial voice indicating the lines of the character in the specific frame, a sound effect corresponding to the onomatopoeia in the specific frame, and an artificial voice indicating a commentary for the specific frame. In this way, the data processing device 12 allows a user who is visually impaired to know that it is time to switch the frame displayed on the display 40A.

[0113] Furthermore, in the data processing device 12, when manga A has been animated, processor 28 extracts from the voice actor database a specific virtual voice actor corresponding to the specific voice actor who voiced the character. Then, when an image of a specific frame is displayed on display 40A, processor 28 causes speaker 40B to output the lines of the character in the specific frame in an artificial voice uttered by the specific virtual voice actor. This makes it possible for data processing device 12 to reduce the sense of discomfort felt by a user who is familiar with the voice of the character in the animated version when hearing the artificial voice of the character.

[0114] Furthermore, in the data processing device 12, if manga A has not been animated, the processor 28 inputs digital comic A into the data generation model 58 and acquires interpreted character characteristics. The processor 28 then inputs a prompt to the data generation model 58, including the acquired character characteristics and an instruction to identify a voice actor with a voice quality suitable for the characteristics. The processor 28 then extracts a virtual voice actor corresponding to the voice actor output by the data generation model 58 from the voice actor database. When an image of a specific frame is displayed on the display 40A, the processor 28 outputs the lines of the character in the specific frame from the speaker 40B using an artificial voice provided by the virtual voice actor. Thus, according to the data processing device 12, even if manga A has not been animated, an artificial voice capable of reproducing a voice quality suitable for the character's characteristics can be assigned to the character. The virtual voice actor is an example of a "second virtual actor" according to the technology of the present disclosure.

[0115] (others) When there are multiple specific voice actors who voiced the character to be processed in the animation, processor 28 may accept the user's selection of one voice actor from the multiple specific voice actors in the artificial voice determination process. In this case, processor 28 extracts a virtual actor corresponding to the voice actor selected by the user from the voice actor database. Then, when an image of a specific frame is displayed on display 40A, processor 28 outputs the lines of the character in the specific frame from speaker 40B using an artificial voice by the virtual voice actor. In this way, when there are multiple specific voice actors, data processing device 12 can set an artificial voice for the character that suits the user's preferences. The virtual voice actor is an example of a "first virtual actor" according to the technology of the present disclosure.

[0116] In the above embodiment, an electronic comic is used as an example of "electronic content" according to the technology of the present disclosure, but this is not limiting. For example, an example of "electronic content" may be other electronic books that are electronic versions of paper textbooks or picture books containing text and illustrations. Furthermore, the specific processing according to the present embodiment may be performed on product packaging (e.g., candy bags) in addition to books.

[0117] In the above embodiment, animation is an example of "visualization" according to the technology of the present disclosure, but the present disclosure is not limited to this. For example, an example of "visualization" may be a movie or the like.

[0118] In the above embodiment, a voice actor is an example of an "actor" according to the technology of the present disclosure, but is not limited to this. For example, an "actor" may be an actor, an actress, an idol, or the like.

[0119] In the above embodiment, when the output of sound corresponding to the image of a specific frame from speaker 40B has ended, processor 28 uses the sound output function of speaker 40B to output the specific sound. Alternatively or additionally, when the output of sound corresponding to the image of a specific frame from speaker 40B has ended, processor 28 may generate a specific vibration using the vibration function of a vibration unit (not shown) included in smart device 14. The vibration unit is a well-known vibration mechanism such as a motor and weight that is installed in various smartphones. In this case, processor 28 instructs smart device 14 to generate a specific vibration based on the end of the output of sound corresponding to the image of a specific frame from speaker 40B. This causes the vibration unit to generate a specific vibration.

[0120] In the above embodiment, the first specific processing shown in Figures 5 to 9 is performed before the user views the digital comic, but this is not limiting. For example, the first specific processing may be performed in real time in parallel with the second specific processing while the user is viewing the digital comic.

[0121] In the above embodiment, the user may be able to reselect the virtual voice actor who will deliver the character's lines while viewing the digital comic. The reselection can be performed by the user's input via the reception device 38. This allows the user to select a virtual voice actor until the artificial voice matches the user's image, even if the artificial voice output from the speaker 40B differs from the user's image.

[0122] In the above embodiment, it may be possible to select the language to be output from speaker 40B. This selection can be made by user input via reception device 38. This makes it possible to output audio in a language different from the language of the text written in the digital comic.

[0123] In the above embodiment, the text-to-speech software on the smart device 14 generates sound data in which an artificial voice reads the lines of a character in a frame based on the character recognition results of the character's lines and the estimated emotion of the character generated by ChatGPT included in the data generation model 58. However, the method of generating the sound data is not limited to this. For example, the character recognition results of the character's lines and the estimated emotion of the character may be input to EmotiVoice included in the data generation model 58, and the sound data may be generated as an output of the data generation model 58. In this way, when the sound data is generated by the data generation model 58 on the data processing device 12, the processor 28 transmits the generated sound data to the processor 46 before or while the user is viewing the digital comic. Then, when the image of the frame is displayed on the display 40A, the processor 46 instructs the speaker 40B to output the acquired sound data, causing the speaker 40B to output the artificial voice indicated in the sound data.

[0124] In the above embodiment, a prompt including an onomatopoeia interpretation result, which is the output result of ChatGPT included in the data generation model 58, and an instruction to generate a sound effect according to the onomatopoeia interpretation result, is input to an Audiobox included in the data generation model 58 to generate sound data representing the sound effect, but the method of generating the sound data is not limited to this. For example, the sound data does not necessarily have to be generated by the data generation model 58, but may be generated using known software capable of generating predetermined sound effects.

[0125] The data processing system 10 according to the present disclosure has been described above mainly with reference to the functions of the data processing device 12. However, the data processing system 10 is not necessarily implemented on a server. The data processing system 10 may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone or the like. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[0126] In the above embodiment, an example was given in which the specific processing is performed by the computer 22 of one data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed to multiple computers including the computer 22, for example, the computer 22 and the computer 36 of the smart device 14. In this case, the processor 28 and the processor 46 are examples of a "processor" according to the technology of the present disclosure.

[0127] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0128] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0129] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0130] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[0131] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[0132] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0133] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[0134] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0135] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference. [Explanation of symbols]

[0136] 10 Data Processing System 12 Data Processing Device 14 Smart Devices 290 Special Processing Department< / url:>

Claims

1. a processor; The processor: Acquire electronic content that includes text and illustrations, inputting an image of a specific section of the content divided into predetermined sections and a prompt including an instruction to estimate the emotion of a character shown in the image of the specific section into a generative model that generates information according to input data; when an image of the specific section is displayed on a display unit capable of displaying the electronic content, outputting the lines of the character in the specific section from an output unit in an artificial voice generated based on the emotion of the character estimated by the generative model. Data processing device.

2. The processor: If the image of the specific section includes an onomatopoeia, inputting the image of the specific section and a prompt including an instruction to interpret the onomatopoeia into the generative model; when the image of the specific section is displayed on the display unit, outputting from the output unit a sound effect generated based on the interpretation result of the onomatopoeia output by the generative model.

2. The data processing device according to claim 1.

3. The processor: when a predetermined operation is received from a user while the image of the specific section is being displayed on the display unit, an estimated content of the emotion of the character generated by the generative model is output from the output unit in a predetermined artificial voice.

2. The data processing device according to claim 1.

4. The processor: when the output of the sound corresponding to the image of the specific section from the output unit is completed, at least one of the sound output function of the output unit and the vibration function of the vibration unit is used to output at least one of the specific sound and the specific vibration.

2. The data processing device according to claim 1.

5. The processor: When the content is visualized, a specific virtual actor corresponding to a specific actor who voiced the character is extracted from an actor database in which a plurality of virtual actors capable of outputting an artificial voice with a predetermined voice quality is stored; when an image of the specific section is displayed on the display unit, the lines of the character in the specific section are output from the output unit in an artificial voice of the specific virtual actor.

2. The data processing device according to claim 1.

6. The processor: If there are a plurality of the specific actors, accepting a user's selection of one actor from the plurality of the specific actors; extracting a first virtual actor corresponding to the one actor whose selection has been accepted from the actor database; when an image of the specific section is displayed on the display unit, the lines of the character in the specific section are output from the output unit in an artificial voice by the first virtual actor.

6. A data processing device according to claim 5.

7. The processor: If the content is not visualized, input the electronic content into the generative model to obtain interpreted characteristics of the character; inputting a prompt including the acquired characteristics of the character and an instruction to ask for an actor having a voice quality suitable for the characteristics into the generative model; extracting a second virtual actor corresponding to the actor output by the generative model from an actor database storing a plurality of virtual actors capable of outputting an artificial voice with a predetermined voice quality; when an image of the specific section is displayed on the display unit, the lines of the character in the specific section are output from the output unit in an artificial voice by the second virtual actor.

2. The data processing device according to claim 1.

8. Acquire electronic content that includes text and illustrations, inputting an image of a specific section of the content divided into predetermined sections and a prompt including an instruction to estimate the emotion of a character shown in the image of the specific section into a generative model that generates information according to input data; when an image of the specific section is displayed on a display unit capable of displaying the electronic content, outputting the lines of the character in the specific section from an output unit in an artificial voice generated based on the emotion of the character estimated by the generative model. A data processing method in which processing is performed by a computer.

9. Acquire electronic content that includes text and illustrations, inputting an image of a specific section of the content divided into predetermined sections and a prompt including an instruction to estimate the emotion of a character shown in the image of the specific section into a generative model that generates information according to input data; when an image of the specific section is displayed on a display unit capable of displaying the electronic content, outputting the lines of the character in the specific section from an output unit in an artificial voice generated based on the emotion of the character estimated by the generative model. A data processing program that causes a computer to perform processing.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A