Processing apparatus, processing program, and processing method

The processing system allows users to selectively output speech information from specific objects in content, addressing the lack of user-friendly features in existing systems by using identification information to manage voice inputs and outputs, thereby enhancing user experience.

JP2026086875APending Publication Date: 2026-05-26Flect Co., Ltd.
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Flect Co., Ltd.
Filing Date
2026-03-02
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing video distribution systems lack user-friendly features for recipients to selectively output speech information associated with specific objects in content, such as characters, leading to unnecessary noise and user inconvenience.

Method used

A processing system that includes a server and terminal devices, capable of receiving and processing speech information associated with multiple objects, allowing users to select and output speech information from a single object while muting others, using identification information to manage voice inputs and outputs.

Benefits of technology

Enables users to selectively output desired speech information, enhancing user experience by reducing unnecessary noise and improving control over content playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086875000001_ABST
    Figure 2026086875000001_ABST
Patent Text Reader

Abstract

To provide a processing device, processing program, and processing method that are more user-friendly for users such as recipients. [Solution] A processing device comprising at least one processor is provided, wherein the at least one processor is configured to receive speech information input from the sender terminal device via a communication interface, associated with each of a plurality of objects included in content generated in the sender terminal device, to select at least one of the plurality of objects via an input interface, and to output the speech information via an output interface, to output the speech information associated with the selected at least one object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , ,

[0005] , , ,

[0001] The present disclosure relates to a processing device, a processing program, and a processing method for outputting speech information associated with a selected object.

Background Art

[0002] Conventionally, a video distribution system via the Internet has been known. For example, Patent Document 1 describes "a recruitment management unit that notifies a user terminal of recruitment terms including video distribution conditions and acquires a posted video from the user terminal, a video analysis unit that analyzes whether the posted video can be distributed and sets the posted video that can be distributed as a distributed video, and a video distribution management unit that distributes the distributed video."

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Therefore, based on the above technologies, an object of the present disclosure is to provide a processing device, a processing program, and a processing method that are more user-friendly for users such as recipients in various embodiments.

Means for Solving the Problems

[0005] According to one aspect of the present disclosure, a processing device is provided comprising at least one processor, wherein the at least one processor is configured to receive, via a communication interface, speech information input from the sender terminal device in association with each of a plurality of objects included in content generated in the sender terminal device, to select at least one of the plurality of objects via an input interface, and to perform processing for outputting speech information associated with the selected at least one object when outputting the speech information via an output interface.

[0006] According to one aspect of the present disclosure, a processing program is provided which causes a computer having at least one processor to function to receive speech information input from a sender terminal device via a communication interface, associated with each of a plurality of objects included in content generated in a sender terminal device, to select at least one of the plurality of objects via an input interface, and to output the speech information via an output interface, by performing processing to output the speech information associated with the selected at least one object.

[0007] According to one aspect of the present disclosure, a processing method is provided for a computer having at least one processor, the method being performed by the at least one processor, the method comprising: receiving speech information from a sender terminal device via a communication interface, which is input in association with each of a plurality of objects included in content generated in a sender terminal device; selecting at least one object from the plurality of objects via an input interface; and outputting speech information associated with the selected at least one object when outputting the speech information via an output interface. [Effects of the Invention]

[0008] This disclosure makes it possible to provide a more user-friendly processing device, processing program, and processing method for users such as recipients.

[0009] The effects described above are merely illustrative for the sake of explanation and are not limiting. In addition to, or in lieu of, any other effects described herein or that would be obvious to those skilled in the art may be achieved. [Brief explanation of the drawing]

[0010] [Figure 1A] Figure 1A is a diagram showing an overview of the processing according to the processing system 1 according to the embodiment of this disclosure. [Figure 1B] Figure 1B is a block diagram showing the configuration of a processing system 1 according to one embodiment of the present disclosure. [Figure 2A] Figure 2A is a block diagram showing the configuration of a server device 100 according to one embodiment of the present disclosure. [Figure 2B] Figure 2B is a block diagram showing the configuration of a terminal device 200 according to one embodiment of the present disclosure. [Figure 3] Figure 3 is a schematic diagram showing information transmitted from the sender terminal device 200-1 as transmission information according to one embodiment of the present disclosure. [Figure 4] Figure 4 shows a processing sequence executed by a processing system 1 according to one embodiment of the present disclosure. [Figure 5] Figure 5 shows the processing flow performed in a receiver terminal device 200-2 according to one embodiment of the present disclosure. [Figure 6] Figure 6 shows an example of a screen output in a sender terminal device 200-1 according to one embodiment of the present disclosure. [Figure 7] Figure 7 shows an example of a screen output in a receiver terminal device 200-2 according to one embodiment of the present disclosure. [Figure 8A]Figure 8A is a diagram showing a processing sequence executed by a processing system 1 according to one embodiment of the present disclosure. [Figure 8B] Figure 8B is a diagram showing a processing sequence executed by a processing system 1 according to one embodiment of the present disclosure. [Figure 9] Figure 9 is a diagram showing an overview of the processing according to the processing system 1 according to the embodiment of this disclosure. [Modes for carrying out the invention]

[0011] 1. Overview of Processing System 1 The processing system 1 described herein is used to output speech information associated with an object desired by the receiver for content transmitted by the sender. For example, the processing system 1 is used to output only the audio associated with a character object when the receiver selects one of the character objects appearing in video content transmitted by the sender.

[0012] Here, Figure 1A is a diagram showing an overview of the processing related to the processing system 1 according to an embodiment of this disclosure. Specifically, Figure 1A shows an example of the processing in the distribution of video content performed using the processing system 1. According to Figure 1A, the user, who is the sender, uses an available sender terminal device to transmit video content in which the objects of character A and character B appear to the receiver terminal device of the user, who is the receiver. For example, the sender himself inputs both the voice information of character A (voice A) and the voice information of character B (voice B) into the video content (typically, it is assumed that the sender plays both character A and character B in the video content).

[0013] Then, the user who is the recipient uses an available recipient terminal device to receive and play video content from the sender terminal device via the server device. By the way, the recipient may have a need to output only the voice of either character A or character B from the video content being played, conversely, to mute the other, according to, for example, their own preferences or their own situation. At this time, if it is only a simple change in the volume of the voice or the voice setting of the playback application, it is only possible to mute both voices of the sender terminal device, that is, both voice A of character A and voice B of character B, or to continue outputting both. However, according to the processing system 1, since identification information for identifying each voice information has been attached in advance to the voice information of voice A of character A and the voice information of voice B of character B, it is possible to output only the desired voice of the recipient and restrict the output of the other, that is, to mute it. In the example of FIG. 1A, only voice A is output, and the output of voice B is restricted, that is, muted.

[0014] Such a processing system 1 is typically used in video content in which character A or character B appears, but it can also be used in other video content such as video conferences and telephone conferences. In such cases as well, in the same way as above, it is possible to restrict the output of either voice by designating the character or identification information of the user participating in the video conference or telephone conference.

[0015] Thus, in the processing system 1, in the sender terminal device, utterance information (e.g., voice information of voice A and voice information of voice B) is input in association with each of a plurality of objects (e.g., character A and character B) included in the content (e.g., video content). On the other hand, in the receiver terminal device, at least one object (e.g., character A) among the plurality of objects is selected. Then, in the receiver terminal device, output of the utterance information (e.g., voice information of voice A) associated with the selected at least one object (e.g., character A) is permitted, and output of the utterance information (e.g., voice information of voice B) associated with an object other than the said object (e.g., character B) is restricted.

[0016] Note that in the present disclosure, "sender" and "receiver" are merely names given to distinguish the person who transmits the content from the person who receives the content. That is, even if described as a sender, when receiving content from another person, it can become a receiver, and even if described as a receiver, when transmitting content to another person, it can become a sender. Also, the sender and the receiver are not limited to only individuals, and may be organizations such as companies or groups. Further, mainly described is the case where the sender himself / herself generates the content, but the sender and the person who generates the content may be different. In this case, even if the person who generates the content only generates the content and does not perform the generation of the content, when the generated content is transmitted by any person, it is included in the sender.

[0017] Also, in the present disclosure, "sender terminal device" and "receiver terminal device" are merely names given to distinguish the terminal device that transmits the content from the terminal device that receives the content. That is, even if described as a sender terminal device, when receiving content from another terminal device, it can become a receiver terminal device, and even if described as a receiver terminal device, when transmitting content to another terminal device, it can become a sender terminal device.

[0018] In this disclosure, "content" means a set of electronic information transmitted and received via a communication network. Examples of such content include video content, music content, game content, publication content, chat content, social networking service (SNS) content, web content, and combinations thereof. Among these, processing system 1 is preferably used for video content that includes at least image information in which multiple character objects are characters and audio information associated with each character object. In this disclosure, video content includes not only content distributed through video content distribution sites, but also, for example, video conferencing content (including cases where only audio is transmitted and received with the camera function turned off), telephone conferencing content, and electronic advertising content such as digital signage. Furthermore, unless otherwise specified below, the example of content will be described as video content, but of course, content is not limited to video content.

[0019] In this disclosure, “object” means data contained within content or means for manipulating and inputting it. Examples of such objects include character objects, structural objects, decorative objects, text objects, image objects, GUI objects, and combinations thereof. Among these, processing system 1 is preferably used for character objects that are included as characters in video content (for example, character A and character B in Figure 1A). In the following, unless otherwise specified, the example of an object will be described as a character object, but of course, objects are not limited to character objects.

[0020] In this disclosure, "processing device" means any of the devices that constitute the processing system 1, and may be a server device, a sender terminal device, or a receiver terminal device. Furthermore, the processing device is not limited to any single device, but may be a combination of multiple devices that can distribute the processing performed by the processing device. "Processing program" and "processing method" mean the program and method executed in the processing device.

[0021] 2. Configuration of Processing System 1 Figure 1B is a block diagram showing the configuration of a processing system 1 according to one embodiment of the present disclosure. According to Figure 1B, the processing system 1 includes a server device 100 for processing content (e.g., video content), a sender terminal device 200-1 for transmitting content, and a receiver terminal device 200-2 for receiving content, which are connected to each other via a communication network.

[0022] In Figure 1B, the sender terminal device 200-1 and the receiver terminal device 200-2 are shown as single devices, but naturally, each may include multiple devices.

[0023] Furthermore, although Figure 1B shows a single server device 100, multiple server devices and other devices may be combined to distribute processing and storage. In this case, the server device 100 may include a combination of multiple server devices and other devices.

[0024] 3. Configuration of Server Device 100 Figure 2A is a block diagram showing the configuration of a server device 100 according to one embodiment of the present disclosure. The server device 100 does not need to have all of the components shown in Figure 2A; it is possible to omit some components or add other components. The server device 100 does not need to have all the components shown in Figure 2A in a single enclosure; it is possible to distribute each component and processing of the server device 100 across multiple devices.

[0025] According to Figure 2A, the server device 100 includes a processor 111 consisting of a CPU, RAM, ROM, and memory 112 including non-volatile memory and an HDD, and a communication interface 113. These components are electrically connected to each other via control lines and data lines.

[0026] The processor 111 consists of a CPU (microcomputer) and functions as a control unit for controlling other connected components based on various programs stored in the memory 112. The processor 111 reads and executes programs from the memory 112 for running the application related to this disclosure and programs for running the OS. Specifically, the processor 111 executes processes such as "receiving speech information input from the sender terminal device 200-1 via the communication interface 113, associated with each of a plurality of objects included in the content generated in the sender terminal device 200-1," and "transmitting the received content to the receiver terminal device 200-2 via the communication interface 113," based on programs stored in the memory 112. The processor 111 mainly consists of one or more CPUs, but may be combined with a GPU, FPGA, etc. as appropriate.

[0027] Memory 112 includes RAM, ROM, non-volatile memory, and HDD, and functions as a storage unit. ROM stores instruction commands as programs for executing the applications and OS related to this disclosure. Such programs are loaded and executed by processor 111. RAM is used to write and read data while the programs stored in ROM are being processed by processor 111. Non-volatile memory is memory on which data is written and read as a result of program execution, and data written therein is retained even after the program execution has finished. Specifically, memory 112 stores programs for processor 111 to perform the above-mentioned processes.

[0028] The communication interface 113 functions as a communication unit that transmits and receives information with other devices such as the remotely installed transmitter terminal device 200-1 and receiver terminal device 200-2 via a communication processing circuit and antenna. The communication processing circuit performs processing to transmit and receive information such as programs and various information used in the processing system 1 as processing progresses. The communication processing circuit processes based on a wideband wireless communication method such as the LTE method, but it can also process based on a narrowband wireless communication method such as wireless LAN such as IEEE 802.11 or Bluetooth®, or a contactless wireless communication method. In addition, it is possible to use wired communication instead of or in addition to wireless communication.

[0029] 4. Configuration of terminal device 200 Figure 2B is a block diagram showing the configuration of a terminal device 200 according to one embodiment of the present disclosure. The terminal device 200 does not need to have all of the components shown in Figure 2B; it is possible to omit some components or add other components. Furthermore, although the terminal device 200 is used as a sender terminal device 200-1 or a receiver terminal device 200-2, the two do not need to have the same configuration; each terminal device may have a different configuration.

[0030] According to Figure 2B, the terminal device 200 includes a processor 211 consisting of a CPU, a memory 212 including RAM, ROM, non-volatile memory, and an HDD, a communication interface 213, an input interface 214, and an output interface 215. These components are electrically connected to each other via control lines and data lines.

[0031] The processor 211 consists of a CPU (microcomputer) and functions as a control unit for controlling other connected components based on various programs stored in the memory 212. The processor 211 reads and executes programs for running the application and the OS from the memory 212. The processor 211 is mainly composed of one or more CPUs, but may be combined with a GPU, FPGA, etc. as appropriate.

[0032] When the processor 111 functions as a sender terminal device 200-1, it executes the following processes based on a program stored in memory 212: "a process of receiving operation input from the sender via the input interface 214 and starting an application program for generating content," "a process of receiving operation input from the sender via the input interface 214 and selecting one of several objects included in the content based on identification information entered to identify that one object," "a process of inputting speech information associated with the selected one object via the input interface 214," and "a process of transmitting content including image information and speech information to the server device 100 via the communication interface 213."

[0033] Furthermore, when the processor 111 functions as a receiver terminal device 200-2, it executes the following processes based on a program stored in memory 212: "receiving operation input from the receiver via the input interface 214, starting an application program for outputting content, and selecting the desired content"; "receiving content from the sender terminal device 200-1 via the communication interface 213, including speech information input in association with each of the multiple objects included in the content generated at the sender terminal device 200-1"; "outputting the selected content via the output interface 215"; "selecting at least one object from the multiple objects included in the content via the input interface 214"; and "outputting speech information associated with the selected at least one object when outputting speech information included in the content via the output interface 215".

[0034] Memory 212 includes RAM, ROM, or non-volatile memory and functions as a storage unit. ROM stores instruction programs for executing applications and operating systems related to this disclosure. Such programs are loaded and executed by the processor 211. RAM is used to write and read data while the programs stored in ROM are being processed by the processor 211. Non-volatile memory is memory on which data is written and read as a result of program execution, and the data written thereto is retained even after the program execution has finished. Specifically, memory 212 stores programs for the processor 211 to perform the above-mentioned processing, etc.

[0035] The communication interface 213 functions as a communication unit that transmits and receives information with the electrically connected server device 100 and other terminal devices 200 via the communication processing circuit. The communication processing circuit performs processing to transmit and receive information such as programs and various other information used in the processing system 1 as processing progresses. The communication processing circuit processes based on a wideband wireless communication method such as the LTE method, but it can also process based on a narrowband wireless communication method such as wireless LAN such as IEEE 802.11 or Bluetooth®, or a contactless wireless communication method. In addition, it is possible to use wired communication instead of or in addition to wireless communication.

[0036] The input interface 214 functions as an input unit that accepts operation input from the sender or receiver to the terminal device 200, as well as input of various types of information by the sender or receiver. Examples of the input interface 214 include various hard keys such as keyboards and mice, a touch panel superimposed on the display of a display device and having an input coordinate system corresponding to the display coordinate system of the display device, as well as a microphone for inputting voice information, which is one type of spoken information, and sensors for sensing the external environment, such as a camera for capturing images. In the case of a touch panel, icons corresponding to the command to be input are displayed on the display, and the user or business operator makes a selection for each icon by performing operation input via the touch panel. The method for detecting operation input by the touch panel may be any method, such as capacitive or resistive touch. The input interface 214 does not always need to be physically provided on the terminal device 200, and may be connected as needed via a wired or wireless network.

[0037] The output interface 215 functions as an output unit for outputting various types of information. An example of the output interface 215 is an interface for connecting to an external device or equipment, such as a display device consisting of a liquid crystal panel, an organic EL display, or a plasma display. However, if the terminal device 200 itself has a display, that display can function as the output interface. Also, if it is connected to a display device or the like via a communication interface 213, the communication interface 213 can function as the output interface 215.

[0038] 6. Examples of Content In this embodiment, as described above, content is generated in the sender terminal device 200-1, and the generated content is output to the receiver terminal device 200-2 via the server device 100. Such content includes multiple objects, and speech information is associated with each object. Such content refers to a set of electronic information transmitted and received via a communication network. Examples include video content, music content, game content, publication content, chat content, SNS content, web content, and combinations thereof. Among these, the processing system 1 is preferably used for video content that includes at least image information in which multiple objects, namely character objects, are characters, and audio information associated with each character object. In the following, unless otherwise specified, the explanation will be given using the case where the content is video content as an example, but of course, the content is not limited to video content, and the processing according to this embodiment can be executed similarly even with other types of content.

[0039] Figure 3 is a schematic diagram showing information transmitted from the sender terminal device 200-1 as transmission information according to one embodiment of the present disclosure. Specifically, Figure 3 is a diagram showing an example of video content that is generated in the sender terminal device 200-1, stored in memory 212, and then transmitted to the server device 100 and stored in the content management table.

[0040] According to Figure 3, video content includes image information and audio information, associated with the content ID information of the video content. "Content ID information" is unique to each video content and is used to identify each video content. This content ID information is generated each time new video content is generated at the sender terminal device 200-1 or each time new video content is received at the server device 100.

[0041] "Image information" refers to image data that constitutes the video content. This image information may be a still image, a moving image, or a combination thereof. Such image information may be captured from real space by a camera provided as one of the input interfaces 214 in the sender terminal device 200-1, or it may be virtually generated by the processing of the processor 211. The image information includes at least multiple objects that are identifiable from one another, and object ID information is assigned to each object. For example, in the example in Figure 3, the image information includes a character object of character A with object ID information "B1" and a character object of character B with object ID information "B2". Examples of such objects include character objects, structural objects, decorative objects, text objects, image objects, GUI objects, and combinations thereof. Among these, the processing system 1 is preferably used for character objects that are included as characters in the video content (for example, character A and character B in Figure 1A). In the following, unless otherwise specified, the example of an object will be a character object, but of course, objects are not limited to character objects.

[0042] "Audio information" is a type of speech information and constitutes audio data that makes up the video content. For example, this audio information is audio data in which the sender's voice is input through a microphone provided as one of the input interfaces 214 in the sender terminal device 200-1. However, other types of audio information may also be, for example, audio data that reproduces the voice of a character object based on text information input via the input interface 214, text data that is a text representation of the sender's voice input via the microphone, text data generated based on text information input via the input interface 214, or other data converted from at least one of these. Such audio information is typically stored in association with the object ID information of each object included in the video content. For example, in the example in Figure 3, audio information A is stored in association with "B1", which is the object ID information of character A, audio information B is stored in association with "B2", which is the object ID information of character B, and BGM audio information is stored as audio information that is not associated with any object ID information.

[0043] In other words, as shown in Figure 3, a video content with content ID information "A1" is shown as an example of content. This video content includes image information that is a video consisting of multiple frames from F1 to Fn and having a length from time t0 to time tn. At least one of the frames of this image information includes, as characters, character A with object ID information "B1" and character B with object ID information "B2" as objects. Here, it is assumed that the sender himself is acting out character A and character B respectively using his own sender terminal device 200-1, as illustrated in Figure 1A. Therefore, this video content includes voice information A of character A, which is input started at time t0 and ended at time t2. Also, this video content includes voice information B of character B, which is input started at time t1 and ended at time t4, as a result of receiving operation input from the sender at time t1. Furthermore, the video content includes voice information A of character A, which was inputted at time t3 and ended at time t6, as the sender's input was received at time t3. The video content also includes voice information B of character B, which was inputted at time t4 and ended at time tn, as the sender's input was received at time t5. In addition, background music audio information is included as audio information not associated with any object during the period from time t1 to time t6. That is, in the example in Figure 3, for example, from time t1 to time t2, from time t3 to time t4, and from time t5 to time t6, voice information A of character A and voice information B of character B are played simultaneously.

[0044] Thus, in the case of video content, the content includes image information and audio information associated with the content ID information. The image information includes each frame (image data) that makes up the video, associated with time (e.g., t0 to tn), frame ID information that identifies each frame (e.g., F1 to Fn), and object ID information that identifies each object contained in at least one of the frames of the image information, all synchronized with time (e.g., t0 to tn). The audio information includes each audio data, associated with time (e.g., t0 to tn), and object ID information associated with each audio data (although there may be no associated object ID information).

[0045] The video content shown in Figure 3 is merely one example of content, as described above. Therefore, even when using video content, it is not necessary to include all of the various types of information exemplified above, and additional information may be included.

[0046] Furthermore, as mentioned above, audio information is a type of speech information, and speech information can be any information that can be associated with the object ID information of an object and reproduce the information entered by the sender. In addition to audio information, it can also be text information or image information, for example.

[0047] 7. Processing sequence executed by processing system 1 Figure 4 is a diagram showing a processing sequence executed in a processing system 1 according to one embodiment of the present disclosure. Specifically, Figure 4 is a diagram showing a series of processing sequences from the generation of content in the sender terminal device 200-1 to the output of the generated content in the receiver terminal device 200-2 via the server device 100. Processing in each device is performed by a processor processing a program stored in the memory of each device.

[0048] (A) Processing related to content generation As shown in Figure 4, first, the content generation process is mainly performed in the sender terminal device 200-1. The processor 211 of the sender terminal device 200-1 receives operation input from the sender via the input interface 214, reads the application program for content generation from memory 212, and starts the application program (S11). Once the application program is started, the processor 211 starts storing the image data, which is frame F1, as image information of the content, as shown in Figure 3. At this time, when the processor 211 receives operation input from the sender via the input interface 214 to request input of audio information associated with an object included in the image information of the content (S12), it outputs the audio input screen via the output interface 215.

[0049] Here, Figure 6 shows an example of a screen output in a sender terminal device 200-1 according to one embodiment of the present disclosure. Specifically, Figure 6 shows an example of a voice input screen 10 output in the sender terminal device 200-1 when it receives an operation input from the sender requesting to input voice information in S12 of Figure 4. According to Figure 6, the voice input screen 10 includes an image information display area 11 and an object selection area 12, along with the name of the application program, "Application A". The image information display area 11 outputs image data 13 of frame F1, for example, as the image information currently being recorded. The image data 13 includes an image 14 of character A and an image 15 of character B as character objects.

[0050] Here, character A shown in image 14 is assigned the object ID information "B1", and character B shown in image 15 is assigned the object ID information "B2". This assignment of object ID information is performed, for example, when video recording is being performed by a camera, by the processor 211 executing object detection processing to detect objects included in each frame, and assigning object ID information to each project at the time each object is first detected. Furthermore, this assignment of object ID information is performed, for example, when the sender terminal device 200-1 virtually generates image information in a virtual space, by the processor 211 assigning object ID information to the objects that are drawn at the time of generation. Therefore, in the example of Figure 6, only image 14 of character A and image 15 of character B happen to be included, but if images of other characters are newly included, character ID information for those other characters will be generated.

[0051] The object selection area 12 contains icons for selecting each object corresponding to an object with an object ID among the objects included in the image information. In the example in Figure 6, the object selection area 12 includes a character A icon 16 corresponding to character A and a character B icon 17 corresponding to character B. When the processor 211 receives an operation input from the sender for any of the icons included in the object selection area 12 (for example, either character A icon 16 or character B icon 17) via the input interface 214, it selects the character corresponding to the icon for which the operation input was made. In the example in Figure 6, character A icon 16 is displayed in a way that allows it to be identified from the other icons, indicating that the object ID information of character A (i.e., "B1") has been selected as the object ID information of the character to which the incoming audio information will be associated.

[0052] Returning to Figure 4, as shown in Figure 6, when the object ID information of a character for which voice information input is desired is selected, the processor 211 of the sender terminal device 200-1 accepts the input of voice information via the input interface 214 in association with the object ID information (S13). Specifically, the processor 211 stores the voice data spoken by the sender from the microphone, which is one of the input interfaces 214, as voice information in the memory 212, in synchronization with each frame of image information associated with the time (e.g., T0) during which the voice information is input.

[0053] The processor 211 of the sender terminal device 200-1 repeatedly performs the recording of image information, the selection of object ID information to associate with audio information, and the input of audio information as described in S11 to S13, thereby generating content with content ID information "A1", as illustrated in Figure 3. When the content generation is complete, the processor 211 stores the generated content (T11) in memory 212 in association with the content ID information and transmits the generated content to the server device 100 via the communication interface 213.

[0054] When the processor 111 of the server device 100 receives content from the sender terminal device 200-1, it stores the received content in the content management table (not shown) of memory 112, associating it with the content ID information (S14). Specifically, the processor 111 of the server device 100 stores each piece of information (image information, object ID information, and audio information, etc.) included in the content shown in Figure 3, associating it with the content ID information in the content management table (not shown) of memory 112. With this, the processing related to content generation is completed.

[0055] In Figure 4, the processor 211 of the sender terminal device 200-1 sent the content to the server device 100 when the content generation was completed. However, the content may be divided and sent each time a predetermined number of frames or amount of data is generated.

[0056] Furthermore, although Figure 4 describes the process assuming that audio information is input while image information is being recorded, the processor 211 of the sender terminal device 200-1 may first generate the image information and then input the audio information in synchronization with each frame of the image information. For example, in the example shown in Figure 3, audio information A and audio information B are input simultaneously from time t1 to time t2, but by inputting each audio piece in synchronization with each frame after the image information has been generated, the same sender can perform both character A and character B.

[0057] Furthermore, although not specifically illustrated in Figure 4, as shown in Figure 3, it is also possible to input audio information that is not associated with the object ID information of a specific object (for example, background music audio information).

[0058] (B) Processing related to content output Next, as shown in Figure 4, the processing related to content output is mainly performed in the receiver terminal device 200-2. This processing involves, for example, if the content is video content, selecting the desired video content in the receiver terminal device 200-2 and playing the video content. The processor 211 of the receiver terminal device 200-2 receives operation input from the receiver via the input interface 214, reads an application program for content output from memory 212, and starts the application program (S21). Once the application program is started, the processor 211 outputs a content selection screen via the output interface 215, which displays a list of thumbnail images for selecting one or more video content that can be output in the receiver terminal device 200-2. The processor 211 then receives operation input from the receiver via the input interface 214 to select the thumbnail image of the desired content from the list on the content selection screen and selects the content to be output (S22). The processor 211 sends a content request (T21) to the server device 100 via the communication interface 213, along with content ID information associated with the selected content (for example, content ID information is "A1"), requesting that the content be transmitted.

[0059] When the processor 111 of the server device 100 receives a content request from the receiver terminal device 200-2 via the communication interface 113, it refers to the content management table based on the content ID information (e.g., A1) received along with the request and reads the content (e.g., the information illustrated in Figure 3) (S23). The processor 111 then sends the read content (T22) to the receiver terminal device 200-2 that sent the content request via the communication interface 113.

[0060] When the processor 211 of the receiver terminal device 200-2 receives content via the communication interface 213, it outputs the received content via the output interface 215 (S24). Here, the processor 211 of the receiver terminal device 200-2 can select the audio information to be output via the receiver's operation input if object ID information is associated with the frame constituting the currently outputting image information in the received content via the input interface 214. That is, when the processor 211 receives operation input from the sender requesting the input of audio information associated with an object (S25), it selects the audio information to be output via the output interface 215 and executes a process to change the output audio information (S26). Details of the series of processes related to S24 to S26 will be described later in Figure 5.

[0061] Then, the receiver terminal device 200-2 repeatedly outputs content, selects an object, and modifies the audio information to be output according to the selection in S24-S26, and when time tn is reached, it terminates the output of content. With this, the process related to content output is terminated.

[0062] In Figure 4, the processor 211 of the receiver terminal device 200-2 receives the content from the server device 100 as a single block of data. However, it may also receive the content at predetermined frame counts or data volumes and output it sequentially.

[0063] 8. Processing flow of receiver terminal device 200-2 Figure 5 is a diagram showing the processing flow executed in a server device 100 according to one embodiment of the present disclosure. Specifically, it is a diagram showing the processing flow related to content output performed by the receiver terminal device 200-2 in S24 to S26 of Figure 4. This processing flow is mainly performed by the receiver terminal device 200-2 reading and executing a program stored in memory 212.

[0064] As shown in Figure 5, the processor 211 receives desired content (for example, content with content ID information A1 as shown in Figure 3) from the server device 100 via the communication interface 213 (S111). Upon receiving the content, the processor 211 outputs the received content via the output interface 215. Specifically, the processor 211 outputs the image information contained in the received content sequentially from frame F1 via the display, which is one of the output interfaces 215. The processor 211 also outputs the audio information contained in the received content via the speaker, which is another output interface 215, in synchronization with the frames of the image information being output. In the example in Figure 3, the audio information A of character A is output from time t0, and at time t1, in addition to the audio information A, the audio information B of character B and the BGM audio information are output.

[0065] The processor 211, via the input interface 214, can accept operation input from the receiver via the input interface 214 and select the audio information to be output if object ID information is associated with the frame that constitutes the image information currently being output to the received content. Therefore, the processor 211 determines whether or not an object has been selected by accepting the operation input (S113).

[0066] Here, Figure 7 shows an example of a screen output in a receiver terminal device 200-2 according to one embodiment of the present disclosure. Specifically, Figure 7 shows an example of a content output screen 20 when the receiver terminal device 200-2 receives an operation input from a receiver to select the audio information to be output in S112 to S113 of Figure 5. According to Figure 7, the content output screen 20 includes an image information display area 21 and an object selection area 22, along with the name of the application program, "Application B". The image information display area 21 outputs, for example, image data 23 of frame F3, as the image information currently being output. The image data 23 includes an image 24 of character A and an image 25 of character B as character objects. Character A shown in image 24 is assigned the object ID information "B1", and character B shown in image 25 is assigned the object ID information "B2".

[0067] The object selection area 22 contains icons for selecting each object, i.e., character, corresponding to the object ID information associated with the audio information that is synchronized with the frame of the currently output image information. For example, taking the content output screen 20 at any point between time t1 and time t2 in Figure 3 as an example, both audio information A for character A and audio information B for character B are output at that time. Therefore, the object selection area 22 contains a character A icon 26 corresponding to character A and a character B icon 27 corresponding to character B. When the processor 211 receives an operation input from the receiver for any of the icons in the object selection area 22 (for example, either character A icon 26 or character B icon 27) via the input interface 214, it selects the character corresponding to the icon to which the operation input was made. In the example in Figure 7, the character A icon 26 is displayed in a identifiable manner relative to the other icons, indicating that the object ID information of character A (i.e., "B1") has been selected as the object ID information of the character to which the incoming audio information will be associated.

[0068] In the example shown in Figure 7, the object selection area 22 includes an icon for selecting the audio information to be output, corresponding to the character included in the currently outputting frame. However, this is not limited to this example; for any character appearing in at least one frame throughout the entire content, an icon for selecting audio information may always be included, allowing for the selection of audio information even when no audio information is being output.

[0069] Furthermore, Figures 5 (S113) and 7 illustrate the case where an object is selected by receiving operation input from the receiver via the input interface 214. However, instead of this, or in addition to this, it is also possible to select an object using a sensor such as a microphone or camera as the input interface 214. For example, the processor 211 uses the camera to recognize the attributes of the receiver using the receiver terminal device 200-2 (e.g., age, gender, etc.). The processor 211 then selects an object to output sound based on the recognition result. For example, a digital signage terminal device is prepared as the receiver terminal device 200-2, and the camera mounted on the terminal device recognizes the attributes of the user (receiver) who is viewing the display of the terminal device. If the user (receiver) is recognized as a "child," sounds other than those of child-oriented objects (e.g., animal characters) are muted, and if the user is recognized as an "adult," sounds other than those of adult-oriented objects (e.g., human characters) are muted. In this way, by using a sensor such as a camera as the input interface 214, it is possible to realize a wider variety of selection methods.

[0070] Returning to Figure 5, as shown in Figure 7, when the object ID information of a character from which audio information output is desired is selected, the processor 211 outputs only the audio information of the selected character and restricts (for example, mutes) the output of audio information of other characters (S114). That is, between time t1 and time t2 in Figure 3, when the object ID information of character A is selected, the processor 211 restricts the output from the output interface 215 (for example, speaker) of character B's audio information B so that only character A's audio information A is output. On the other hand, if no object is selected in S113, the process related to S114 is skipped.

[0071] Processor 211 continuously repeats the processes related to S112 to S114 until it has finished outputting a series of content from time t0 to tn. With this, this processing flow is terminated.

[0072] In this explanation, we are describing a case where only the voice information A for character A and the voice information B for character B are included in the content. Therefore, when character A is selected, the output of voice information B is restricted, and only voice information A is output. However, in Figure 7, the output of the voice information of the selected character may be restricted, while the voice information of the unselected character may be output without restriction.

[0073] Furthermore, if the content contains three or more audio pieces, (1) Output only the voice information of the selected character, and restrict the output of voice information for all remaining characters. (2) Restrict the output of the voice information of the selected character, and output the voice information of all remaining characters. (3) Output the voice information of the selected multiple characters and restrict the output of voice information for all remaining characters. (4) Restrict the output of voice information for selected characters, and output the voice information for all remaining characters. Audio information can be output in various combinations, such as those mentioned above.

[0074] Furthermore, while the above example uses "muting" as a method for restricting the output of audio information, various other restriction methods can be employed, such as changing the volume when outputting (for example, lowering it), or outputting subtitle text information simultaneously with normally outputted audio information, but not outputting subtitles for restricted audio information.

[0075] In this embodiment, it is possible to provide a processing device, processing program, and processing method that are more user-friendly for users such as receivers. In particular, when the output content includes multiple speech information (e.g., voice information), it is possible for the receiver to select which speech information (e.g., voice information) to output. For example, conventionally, if it was not possible to output voice information associated with certain objects, the output was limited by controlling the volume using a volume button on the receiver terminal device 200-2, etc. Consequently, the output of all voice information was limited. However, in this embodiment, it is possible to selectively output only the voice information associated with the object desired by the receiver at the timing desired by the receiver, or to selectively limit the output.

[0076] 9. Variations The following shows modified examples of the above embodiment shown in Figures 1 to 7. Note that the following modified examples and the embodiment shown in Figures 1 to 7 can be combined and implemented. Furthermore, except for points specifically mentioned below, the process can be carried out in the same manner as described in the embodiment shown in Figures 1 to 7.

[0077] (A) Variation 1 of the selection of audio information In the above, as shown in Figure 4, etc., the case was described in which the audio information associated with the character selected in the receiver terminal device 200-2 is selected by the processor 211 of the receiver terminal device 200-2, and the output of audio information for other characters is restricted. However, instead, the audio information associated with the character selected in the receiver terminal device 200-2 may be reorganized by the processor 111 of the server device 100, and the output of audio information for other characters may be restricted.

[0078] Figure 8A is a diagram showing a processing sequence executed in a processing system 1 according to one embodiment of the present disclosure. Specifically, Figure 8A is a diagram showing a processing sequence when the processing related to the selection of speech information, which is one of the speech information, is performed by the processor 111 of the server device 100. The processing in each device is executed by the processor processing a program stored in the memory of each device.

[0079] Note that the processes related to content generation in S31-S34 are the same as the processes related to content generation in S11-S14 shown in Figure 4, so their explanation will be omitted.

[0080] Furthermore, regarding the content output process, the processes S41 to S45, which involve receiving the content output, starting the application program, outputting the desired content via the output interface, and selecting the audio information associated with the desired object, are the same as the processes S21 to S25 in the content output process shown in Figure 4, so their explanation will be omitted.

[0081] When the object ID information of a character from which voice information output is desired is selected using the method shown in Figure 7, the processor 211 of the receiver terminal device 200-2 transmits object selection information (T43), which includes the content ID information of the content to be output and the selected object ID information, to the server device 100 via the communication interface 213.

[0082] When the processor 111 of the server device 100 receives object selection information, it reads the content associated with the content ID information from memory 112 and performs a process to reorganize the content (S46). Specifically, the processor 111 keeps the audio information associated with the selected object ID information (audio information A in the example of Figure 7) as is, and deletes the other audio information that was not selected (audio information B in the example of Figure 7) from the content. After reorganizing the content through the above process, the processor 111 stores the reorganized content in memory 112 and transmits the content (T44) to the receiver terminal device 200-2 that sent the object selection information via the communication interface 113.

[0083] When the receiver terminal device 200-2 receives the reorganized content via the communication interface 213, it outputs the received content via the output interface 215, similar to S44. At this time, the audio information of the content does not include the audio information B of character B. Therefore, the processor 111 of the receiver terminal device 200-2 outputs only the audio information A of character A via the output interface 215, without outputting the audio information B of character B.

[0084] Furthermore, for example, in the object selection area 22 of the content output screen 20 output in Figure 7, icons corresponding to objects (characters) included in at least some of the frames of the image information will always be displayed. This ensures that even if the output of audio information B is restricted, if the recipient wishes to output audio information B again, they will be able to select audio information B.

[0085] As shown in Figure 8A, selective output of audio information is possible, similar to the embodiments in Figures 1 to 7.

[0086] (B) Modification example 2 of the selection of audio information In the above, as shown in Figure 4, etc., the case in which the audio information associated with the character selected in the receiver terminal device 200-2 is selected by the processor 211 of the receiver terminal device 200-2, and the output of audio information for other characters is restricted, has been described. However, instead, the audio information associated with the character selected in the receiver terminal device 200-2 may be selected by the processor 211 of the sender terminal device 200-1, and the output of audio information for other characters may be restricted.

[0087] Figure 8B is a diagram showing a processing sequence executed in a processing system 1 according to one embodiment of the present disclosure. Specifically, Figure 8B is a diagram showing a processing sequence when the processing related to the selection of voice information, which is one of the utterance information, is performed by the processor 211 of the sender terminal device 200-1. The processing in each device is executed by the processor processing a program stored in the memory of each device.

[0088] Except for the fact that the content is streamed in fixed data chunks, the processing in steps S61 to S64 related to content generation is the same as the processing in steps S11 to S14 related to content generation shown in Figure 4, so its explanation will be omitted.

[0089] Furthermore, except for the fact that the content is streamed in fixed data units, the processing related to content output, specifically steps S71 to S75, which involve receiving the content output, starting the application program, outputting the desired content via the output interface, and selecting the audio information associated with the desired object, is the same as the processing related to content output S21 to S25 shown in Figure 4, so its explanation will be omitted.

[0090] When the object ID information of a character from which voice information output is desired is selected using the method shown in Figure 7, the processor 211 of the receiver terminal device 200-2 transmits object selection information (T43), which includes the content ID information of the content to be output and the selected object ID information, to the server device 100 via the communication interface 213.

[0091] When the processor 111 of the server device 100 receives object selection information (T73), it identifies the sender terminal device 200-1, which is the sender of the content, based on the content ID information (S76). Then, the processor 111 transmits object selection information (T74) to the identified sender terminal device 200-1 via the communication interface 113.

[0092] When the processor 211 of the sender terminal device 200-1 receives object selection information via the communication interface 213, it selectively inputs audio information (S77). Specifically, in the sender terminal device 200-1, image information and audio information are input and distributed in real time. The processor 211 accepts audio information input associated with object ID information through the processes shown in S62 and S63 (i.e., the processes shown in S22 and S23 in Figure 4). The processor 211 then refers to the object ID information received via the object selection information, and if audio information associated with the same object ID information as the received object ID information is input, it stores the audio information in synchronization with the image information. On the other hand, the processor 211 accepts audio information associated with object ID information different from the received object ID information, but does not include it in the content to be transmitted. In other words, the processor 211 generates content that includes only audio information associated with the object ID information of the character selected by the receiver, and does not include audio information associated with the object ID information of other characters.

[0093] The processor 211 of the sender terminal device 200-1 transmits the content (T75) generated as described above, along with content ID information, to the server device 100 via the communication interface 213. When the processor 111 of the server device 100 receives the content via the communication interface 113, it stores it in the content management table in association with the content ID information, and also transmits the received content (T76) via the communication interface 113 to the receiver terminal device 200-2 that sent the object selection information.

[0094] When the processor 211 of the receiver terminal device 200-2 receives content via the communication interface 213, it outputs the received content via the output interface 215. At this time, the received content includes only the audio information associated with the object ID information of the selected character, as described above, and does not include audio information associated with the object ID information of other characters. In other words, the output of audio information associated with object ID information of characters other than the one selected by the receiver is restricted because its transmission is restricted.

[0095] As described above, the example shown in Figure 8B also enables selective output of audio information, similar to the embodiments shown in Figures 1 to 7.

[0096] (C) Modifications relating to restricted audio information Figures 1 to 8B illustrate the case where the content contains only voice information A for character A and voice information B for character B. Therefore, when character A is selected, the output of voice information B is restricted, and only voice information A is output. However, it is also possible to restrict the output of the voice information of the selected character while outputting the voice information of the other character without restriction.

[0097] Furthermore, if the content contains three or more audio pieces, (1) Output only the voice information of the selected character, and restrict the output of voice information for all remaining characters. (2) Restrict the output of the voice information of the selected character, and output the voice information of all remaining characters. (3) Output the voice information of the selected multiple characters and restrict the output of voice information for all remaining characters. (4) Restrict the output of voice information for selected characters, and output the voice information for all remaining characters. Audio information can be output in various combinations, such as those mentioned above.

[0098] (C) Variation with multiple senders In the examples in Figures 1 to 8B, the case where one sender can perform multiple characters was explained by inputting the voice information of multiple characters in association with object ID information in a single sender terminal device 200-1. However, instead of this, or in addition to this, it is also possible for multiple senders to perform the same character, or for multiple senders to perform multiple characters, by inputting the voice information of multiple characters in association with object ID information in multiple sender terminal devices 200-1.

[0099] Figure 9 is a diagram illustrating an overview of the processing related to the processing system 1 according to an embodiment of this disclosure. Specifically, Figure 9 shows an example of processing in the distribution of video content performed using the processing system 1. According to Figure 9, for the same video content, the sender terminal device of sender A receives voice information A for character A and voice information B for character B, and transmits them to the receiver terminal device of the recipient via the server device. Similarly, the sender terminal device of sender B receives voice information C for character C and voice information D for character D, and transmits them to the receiver terminal device of the recipient via the server device. At this time, sender ID information for identifying sender A or sender A's sender terminal device is associated with voice information A and voice information B. Additionally, sender ID information for identifying sender B or sender B's sender terminal device is associated with voice information C and voice information D. Therefore, when selecting the voice information to be output in the receiver terminal device, it is possible to select sender ID information instead of object ID information. For example, if the recipient terminal device selects the sender ID information of sender A, only voice information A and voice information B will be output, and the output of voice information C and voice information D will be restricted. Similarly, if the recipient terminal device selects the sender ID information of sender B, only voice information C and voice information D will be output, and the output of voice information A and voice information B will be restricted.

[0100] As shown in the example in Figure 9, selective output of audio information is possible, similar to the embodiments in Figures 1 to 8B.

[0101] (D) Variations relating to content, objects, and speech information In the examples in Figures 1 to 8B, video content was used as the content, and the explanation was given using the case where the object is a character object and the speech information is audio information. However, regardless of whether the content is video content or other content, the same processing can be performed with other objects or other speech information. For example, in addition to video content, content can include music content, game content, publication content, chat content, SNS content, web content, and combinations thereof. Similarly, in addition to character objects, objects can include structure objects, decorative objects, text objects, image objects, GUI objects, and combinations thereof. Furthermore, in addition to audio information, speech information can include text information, image information, and combinations thereof.

[0102] For example, when applying chat content as content to the embodiments of this disclosure, the objects include GUI objects in the shape of speech bubbles associated with each sender, and the utterance information includes text information entered by each user as chat. Even in such a case, the recipient can restrict the output (display) of chat (text information) associated with the GUI objects of other senders by selecting the GUI object of the desired sender. This makes it possible to selectively output only specific senders.

[0103] The processes and procedures described herein can be implemented not only by those explicitly described in the embodiments, but also by software, hardware, or a combination thereof. Specifically, the processes and procedures described herein can be implemented by implementing the logic corresponding to the process on a medium such as an integrated circuit, volatile memory, non-volatile memory, magnetic disk, or optical storage. Furthermore, the processes and procedures described herein can be implemented as computer programs and executed by various computers, including processing units and server devices.

[0104] Even if it is stated that the processes and procedures described herein are performed by a single device, software, component, or module, such processes or procedures may be performed by multiple devices, multiple software programs, multiple components, and / or multiple modules. Similarly, even if it is stated that the various types of information described herein are stored in a single memory or storage unit, such information may be distributed and stored in multiple memories within a single device or in multiple memories distributed across multiple devices. Furthermore, the software and hardware elements described herein may be implemented by integrating them into fewer components or by decomposing them into more components. [Explanation of symbols]

[0105] 1. Processing System 100 Server Devices 200 terminal devices 200-1 Sender terminal device 200-2 Receiving terminal device

Claims

[Claim 1] A processing unit comprising at least one processor, The at least one processor is The system receives speech information from the transmitting terminal device via a communication interface, which is input in association with each of the multiple objects included in the content generated in the transmitting terminal device. Select at least one of the multiple objects via the input interface. When outputting the speech information via the output interface, the speech information associated with the selected at least one object is output. A processing unit configured to perform a process for that purpose.