Information processing system, information processing device, and program
The review system uses HRTF/BRTF with HMDs to recreate a reference space for audiovisual content, enabling accurate reproduction and feedback recording for improved content production and review.
Patent Information
- Application Number
- PCT/JP2025/007580
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2025-03-04
- Publication Date
- 2025-10-02
AI Technical Summary
Existing technologies face challenges in producing audiovisual content with 3D CG video and stereophonic sound in a reference environment, as remote production methods are difficult to implement, and it's hard to accurately reproduce spatial sound effects without visual cues, making it difficult to record and reference user feedback during content review.
A review system that combines HRTF/BRTF with HMDs to recreate a reference space for both vision and hearing, allowing users to input feedback through various interfaces, which is then recorded and referenced for content improvement.
Enables accurate reproduction of audiovisual content in a reference space, facilitating effective content review and production by allowing users to provide feedback that can be recorded and utilized for improvements.
Smart Images

Figure JP2025007580_02102025_PF_FP_ABST
Abstract
Description
Information processing system, information processing device, and program
[0001] The present disclosure relates to an information processing system, an information processing device, and a program, and more particularly to an information processing system, an information processing device, and a program that enable more suitable content review and production work.
[0002] The audiovisual content, which combines 3DCG video and stereophonic sound, can give a user who watches the audiovisual content a sufficient sense of realism.
[0003] For example, Patent Literature 1 discloses an audio processing device that converts the acoustic characteristics of audio content of a movie into the acoustic characteristics of a movie theater and provides the converted audio content to a user terminal together with a spherical video including an image of a screen installed in the movie theater and an image of the environment surrounding the screen.
[0004] International Publication No. 2020 / 189263
[0005] When reviewing the above-mentioned audiovisual content while viewing it, a system is required that allows users' feedback to be appropriately recorded and referenced.
[0006] The present disclosure has been made in light of such circumstances, and aims to realize more suitable content review and production work.
[0007] The information processing system of the present disclosure is an information processing system that includes an input accepting unit that accepts input information related to content from a user viewing the content, an acquisition unit that acquires content-related information when the input information is accepted, a recording control unit that records the input information in association with the content-related information, and an output control unit that outputs the input information associated with the content-related information.
[0008] The information processing device of the present disclosure is an information processing device that has an output control unit that outputs input information related to content that is received from a user watching the content and that is associated with content-related information at the time the input information is received, from among input information that is recorded in association with the content-related information.
[0009] The program disclosed herein is a program for causing a computer to execute a process of outputting input information associated with content-related information from among input information about the content received from a user watching the content and content-related information at the time the input information was received, the input information being associated with the content-related information and recorded.
[0010] In the present disclosure, input information related to the content is accepted from a user viewing the content, content-related information is acquired at the time the input information is accepted, the input information is associated with the content-related information and recorded, and the input information associated with the content-related information is output.
[0011] 1 is a diagram illustrating an example configuration of a review system to which the technology according to the present disclosure is applied; FIG. 1 is a diagram illustrating a flow of measuring the transfer characteristics of sound in a reference space; FIG. 2 is a diagram illustrating shooting of a reference space; FIG. 3 is a diagram illustrating an example of a spatial image reproducing a reference space; FIG. 4 is a flowchart illustrating processing of a review application; FIG. 5 is a diagram illustrating an example of an input UI for feedback information; FIG. 6 is a diagram illustrating an example of an input UI for feedback information; FIG. 7 is a diagram illustrating an example of an input UI for feedback information; FIG. 8 is a diagram illustrating an example of a list of feedback information; FIG. 9 is a flowchart illustrating UI processing of the list; FIG. 10 is a diagram illustrating an example of content playback corresponding to a specified record; FIG. 11 is a flowchart illustrating processing to update time information; FIG. 12 is a flowchart illustrating processing for a game scene; FIG. 13 is a flowchart illustrating processing to create a user dictionary; FIG. 14 is a diagram illustrating an example of creating a user dictionary; FIG. 15 is a flowchart illustrating processing to present candidate information; FIG. 16 is a block diagram illustrating an example configuration of hardware of a computer;
[0012] Modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described below in the following order.
[0013] 1. Background 2. Example of review system configuration and reproduction of reference space 3. Example of review application processing and input UI 4. Example of feedback information usage 4-1. Display of feedback information 4-2. Changing content version 4-3. Application to computer games 4-4. Creating a user dictionary 5. Example of computer hardware configuration
[0014] 1. Background Audiovisual content that combines 3D CG video and stereophonic sound can provide a user who watches the audiovisual content with a sufficient sense of realism.
[0015] It is important to produce audiovisual content such as television programs and music videos in a reference environment (also known as a reference space) with suitable conditions to ensure that the production environment does not affect the quality of the content. However, producing in such a reference environment has been difficult using remote production technologies, which have been widely used in recent years.
[0016] Furthermore, although it is necessary to check the spatial sound effects of 3D audio content, it has been difficult to prepare the necessary playback environment.
[0017] On the other hand, HRTF (Head Related Transfer Function) can simulate the acoustic transfer function from the sound source to the eardrum and reproduce the sound at the measurement location using headphones, etc. However, when listening to only the reproduced sound, it is not easy to determine with what accuracy the sound has been reproduced if there is no visual information about the measurement location.
[0018] In response to this, a reference space can be reproduced for both vision and hearing by presenting a view from a specific viewpoint in a reference space using an HMD (Head Mounted Display) that uses binocular stereoscopic vision, while presenting sound at a specific viewpoint in the reference space using HRTF or BRTF (Binaural Room Transfer Function) that has been individually optimized through acoustic measurements.
[0019] However, when reviewing such audiovisual content, a system is required that allows users to appropriately record and refer to their feedback. For example, when a user is viewing content while wearing an HMD, it is difficult to see what is in front of them, making it difficult to record comments or input work instructions when reviewing the content.
[0020] In the following, a review system to which the technology according to the present disclosure is applied will be described in order to realize more suitable content review and production work.
[0021] 2. Configuration Example of Review System and Reproduction of Reference Space> (Configuration Example of Review System) FIG. 1 is a diagram showing a configuration example of a review system that is an embodiment of an information processing system to which the technology according to the present disclosure is applied.
[0022] 1 is used for reviewing and producing content. The content to be reviewed in the review system 1 includes content audio played back using sound transfer characteristics measured in a reference space and content video corresponding to the content audio, but may also include either one of them.
[0023] The review system 1 is configured to include a file storage unit 10 , an HRTF / BRTF superimposition unit 20 , an audio playback device 30 , a playback processing unit 40 , a display controller 50 , a display 60 , an information processing device 100 , and a reference UI / UX 200 .
[0024] The file storage unit 10 stores HRTFs and BRTFs that have been measured in advance in a reference space.
[0025] The HRTF / BRTF superimposing unit 20 superimposes the HRTF and BRTF stored in the file storage unit 10 onto the content audio A1 consisting of multi-channel audio signals corresponding to a plurality of sound sources, and supplies the superimposed audio to the information processing device 100. The HRTF / BRTF superimposing unit 20 may also perform correction processing related to the headphones used when measuring the HRTF and BRTF.
[0026] Furthermore, the content audio A1 may be supplied directly to the information processing device 100 without HRTF or BRTF being superimposed thereon.
[0027] The information processing device 100 performs various acoustic processes on the content audio A1 and supplies the processed audio to the audio playback device 30.
[0028] The audio playback device 30 is configured with headphones, a speaker capable of transaural playback, etc. The audio playback device 30 plays back the content audio A1 from the information processing device 100.
[0029] The playback processing unit 40 executes playback processing of the content video V2 corresponding to the content audio A1 and supplies it to the information processing device 100. The content video V2 may be recorded locally in the review system 1, or may be downloaded or streamed from a server on the cloud or the like.
[0030] In the reproduction processing unit 40, reproduction processing of the content video V2 is executed in synchronization with the content audio A1 based on time information such as a word clock signal and a linear time code obtained from the content audio A1.
[0031] The information processing device 100 acquires a reference space image R3 (hereinafter simply referred to as space image R3) based on the reference space, and supplies a composite image obtained by combining the content image V2 with the space image R3 to the display controller 50.
[0032] The display controller 50 controls the display of the composite image on the display 60. The display controller 50 may be built into the information processing device 100.
[0033] The display 60 is configured, for example, as an HMD, and displays stereoscopic images with binocular parallax to the user. However, the display 60 is not limited to this, and may be configured as a single monitor that allows stereoscopic viewing with the naked eye.
[0034] The information processing device 100 may be configured as a single device such as a PC (Personal Computer), or may be configured as multiple devices such as two or more processors. In the information processing device 100, a predetermined program is executed by the processor constituting the information processing device 100, thereby realizing at least a part of the functional units shown in FIG.
[0035] That is, the information processing device 100 includes an HRTF / BRTF superimposing section 110 , a synthesizing section 120 , an EQ processing section 130 , and a delay adjusting section 140 .
[0036] The HRTF / BRTF superimposing unit 110 superimposes the HRTF or BRTF stored in the file storage unit 10 onto the content audio A1 directly supplied to the information processing device 100, and supplies the superimposed audio to the synthesis unit 120. The HRTF / BRTF superimposing unit 110 may also perform correction processing related to the headphones used when measuring the HRTF or BRTF.
[0037] It should be noted that both the HRTF / BRTF superimposing section 20 and the HRTF / BRTF superimposing section 110 may be provided, or only one of them may be provided.
[0038] When both the HRTF / BRTF superposition unit 20 and the HRTF / BRTF superposition unit 110 are provided, the synthesis unit 120 synthesizes the content audio A1 from the HRTF / BRTF superposition unit 20 and the content audio A1 from the HRTF / BRTF superposition unit 110 and supplies the synthesized audio to the EQ processing unit 130.
[0039] The EQ processing unit 130 performs acoustic processing such as volume adjustment, sound quality adjustment processing, and 2ch mix processing on the content audio A1 from the synthesis unit 120, and supplies the result to the delay adjustment unit 140. The EQ processing unit 130 may also perform correction processing related to the headphones used when measuring the HRTF and BRTF.
[0040] The delay adjustment unit 140 performs delay adjustment processing on the content audio A1 from the EQ processing unit 130 and supplies the resultant audio to the audio playback device 30.
[0041] Furthermore, in the information processing device 100, a predetermined program is executed by the processor, thereby enabling the review application 150 to be executed.
[0042] The review application 150 has, as functional units, a video synthesis unit 151, input reception units 152a, 152b, and 152c, and an acquisition / recording / output control unit 153.
[0043] The video synthesis unit 151 acquires the spatial video R3 and synthesizes the content video V2 from the playback processing unit 40 with the spatial video R3 to generate a synthesized video. The video synthesis unit 151 supplies the synthesized video in which the content video V2 is synthesized with the spatial video R3 to the display controller 50.
[0044] The input accepting units 152a, 152b, and 152c accept input information related to the content to be reviewed by presenting various input UIs to a user who is viewing the content to be reviewed using the audio playback device 30 and the display 60. Here, the input information related to the content is referred to as feedback information F, which includes evaluations and impressions of the content, as well as suggestions and instructions for correction.
[0045] The feedback information F is realized by user input through operation of an input device or the like and user voice through speech, and includes at least one of selection information regarding the content, written information for the content, and voice input information regarding the content.
[0046] The input receiving unit 152a presents pre-registered or externally input template data as options related to content, and receives input of selection information representing the option selected by the user. The selection information includes the content category, the category of evaluation for the content, the details of the evaluation, etc.
[0047] The input receiving unit 152b receives input of entry information based on user input. The entry information includes image data including characters and figures directly entered in the content, text data extracted from the image data, and the like.
[0048] The input receiving unit 152c receives input of voice input information based on the user's voice, which includes text data converted from the user's voice.
[0049] The feedback information received by the input receiving units 152 a , 152 b , and 152 c is supplied to the acquisition, recording, and output control unit 153 .
[0050] The acquisition / recording / output control unit 153 functions as an acquisition unit that acquires content-related information when receiving input of feedback information. Specific examples of content-related information will be described later.
[0051] The acquisition / recording / output control unit 153 also functions as a recording control unit that associates the acquired content-related information with feedback information from the input receiving units 152a, 152b, and 152c and records the information in the feedback DB 160. The feedback DB 160 may be configured to include a DBMS (Database Management System).
[0052] Furthermore, the acquisition / recording / output control unit 153 has a function as an output control unit that outputs feedback information associated with content-related information recorded in the feedback DB 160 to the reference UI / UX 200. The reference UI / UX 200 is configured by, for example, an information presentation device having a touch panel function that is provided outside the information processing device 100. The reference UI / UX 200 may be connected to the information processing device 100 directly or via a network such as the Internet. Furthermore, the reference UI / UX 200 may include an output control unit that controls the output of feedback information associated with content-related information recorded in the feedback DB 160.
[0053] The feedback information output to the reference UI / UX 200 can be referenced not only by users who review content, but also by creators and editors who create the content.
[0054] This configuration allows the user to review content in a reference space that is reproduced visually and aurally.
[0055] (Reproduction of Reference Space) Here, the reproduction of reference space for vision and hearing will be described in detail.
[0056] (1) Measurement of sound transfer characteristics in a reference space The flow of measuring sound transfer characteristics in a reference space will be described with reference to Fig. 2. In the following, an example of measuring HRTF as the sound transfer characteristics will be described, but BRTF can also be measured in the same manner.
[0057] As a first step, the HRTF from the loudspeaker to the ear is measured in a reference space RS. As shown in Fig. 2, the reference space RS is a space equipped with a screen and loudspeakers, such as a movie theater or hall, that is well-suited for the production of audiovisual content.
[0058] First, microphones are inserted into the left and right ear holes of the user and fixed so that each microphone floats in midair within the ear canal.
[0059] Next, measurement signals are reproduced from the speakers installed in the reference space RS, and the measurement signals are recorded by microphones inserted into the ear holes.
[0060] Then, the transfer function from the speaker to the ear is calculated from the recording result and the played-back measurement signal. In the frequency domain, Y is the recording result, X is the measurement signal, and H_Room is the transfer function (HRTF), so Y = H_Room X. In the time domain, y is the recording result, x is the measurement signal, and h_Room is the transfer function (HRIR: Head Related Impulse Response), so y = h_Room x.
[0061] The transfer function H_Room (h_Room) calculated as described above is stored as acoustic transfer function data from the speaker to the ear, including the acoustic characteristics of the room (reference space RS).
[0062] In the second step, the HRTF of the headphones used to play the content audio is measured.
[0063] Again, the same microphone is used as in the first stage. By performing the second stage measurement immediately after the first stage measurement, the timing is kept as close as possible to the first stage measurement, preventing any misalignment of the microphone due to unnecessary user movement.
[0064] Next, measurement signals are played back from the left and right diaphragms of the headphones, and the measurement signals are recorded by microphones inserted into the ear canals.
[0065] Then, the transfer function from the headphone diaphragm to the ear is calculated from the recording result and the played-back measurement signal. In the frequency domain, if the recording result is Y, the measurement signal is X, and the transfer function (HRTF) is H_HP, then Y = H_HP x. In the time domain, if the recording result is y, the measurement signal is x, and the transfer function (HRIR) is h_HP, then y = h_HP x.
[0066] Using the transfer function H_HP calculated as above, an inverse filter inverse_H_HP that cancels the acoustic characteristics of the headphones is calculated and stored as acoustic transfer function data.
[0067] In the third step, the data obtained in the first and second steps is saved as a file.
[0068] That is, the acoustic transfer function data obtained in the first and second stages and all parameters used in the measurements and calculations are saved as a file.
[0069] Using these data, the audio signal calculated from H_Room, X, and inverse_H_HP is played back through the headphones used in the second measurement. This makes it possible to reproduce the acoustic characteristics of the room in which the first measurement was performed with such precision that the sound played back from the room's speakers is indistinguishable from the sound played back from the headphones.
[0070] (2) Output of Multi-Channel Audio Signals An example of output of multi-channel audio signals will be described.
[0071] A DAW (Digital Audio Workstation) can output channel signals such as 2ch, 5.1ch, and 7.1ch.
[0072] In a DAW, an audio signal is created by arranging a large number of recorded data in chronological order and adding them together. The audio signals output from each speaker are called channel signals.
[0073] Software that handles the rendering process of object audio can take 100 or more object tracks as input, render them into channel signals such as 13ch or 16ch, and output them.
[0074] Rendering is a technique in which a sound source and the direction from which the sound from that source should be heard are defined as object data, and the signals distributed to the speakers are adjusted so that the sound is output from the desired direction according to the actual placement of the speakers. Rendering is performed using the object data and the position information of the speakers used for playback, and channel signals are generated as the result of the processing and played from the speakers.
[0075] In the review system 1 described with reference to Figure 1, the channel signal is input as a multi-channel audio signal, but if a rendering processing unit is included in the system, object audio data may also be input.
[0076] (3) Generation of Spatial Image Reproducing Reference Space An example of generation of a spatial image R3 reproducing a reference space will be described with reference to FIG.
[0077] The spatial image R3 is generated as an image that accurately reproduces the field of view seen from the HRTF measurement position in the reference space RS using stereo images with binocular parallax or CG (Computer Graphics) using polygons and point clouds.
[0078] When generating a stereo image with binocular parallax, a camera CAM is placed at the position where acoustic measurements were performed in the reference space RS, aligned with the position of the user's eyes. Then, the camera CAM is moved to capture images (still images or moving images) while taking into account the inter-pupil distance (IPD) and tilt. Taking into account that the IPD and tilt differ from person to person, various patterns are captured while making fine adjustments.
[0079] An image IMG_L captured by a camera CAM placed at the position of the left eye is used as an image for the left eye in the HMD. Similarly, an image IMG_R captured by a camera CAM placed at the position of the right eye is used as an image for the right eye in the HMD.
[0080] (4) Composition of Spatial Image and Content Image Fig. 4 is a diagram showing an example of the spatial image R3 generated as described above and reproducing the reference space RS. In the spatial image R3, the content image V2 is embedded and displayed in the screen area SC corresponding to the screen of the reference space RS based on processing by the image composition unit 151 (Fig. 1).
[0081] The format of the content video V2 may be a video file such as MP4, a video signal transmitted from a PC responsible for video playback, or a video streaming played in a browser, etc. Furthermore, instead of embedding the content video V2 in the screen area SC of the spatial video R3, the reference space RS in which the content video V2 is displayed on the screen may be captured in real time using a stereo camera or the like and distributed.
[0082] In this way, by playing back the content video V2 superimposed on the screen of the CG-displayed reference space RS and the content audio A1 in which the HRTF measured in the reference space RS is superimposed on the corresponding multi-channel audio signal, it is possible to achieve a viewing experience equivalent to that in the reference space RS.
[0083] 3. Examples of Review Application Processing and Input UI The following describes examples of the processing and input UI of the review application 150, which allows a user reviewing content during the viewing experience described above to input feedback information such as suggestions or instructions for correcting aspects of the video or audio that they find concerning.
[0084] (Processing of Review Application) The processing of review application 150 will be described with reference to the flowchart of Fig. 5. The processing of Fig. 5 can be executed when a user is performing a review while viewing content using audio playback device 30 and display 60.
[0085] In step S1, the input receiving units 152a, 152b, and 152c receive feedback information input by users who are viewing content.
[0086] In step S2, the acquisition / recording / output control unit 153 acquires the content-related information when the input of the feedback information is accepted.
[0087] Specifically, the acquisition / recording / output control unit 153 acquires, as content-related information, time information of the content at the time when the input of the feedback information was accepted from the playback processing unit 40 that is executing the playback processing of the content. Furthermore, the acquisition / recording / output control unit 153 may acquire, as content-related information, an extracted scene (capture) of the content at the time when the input of the feedback information was accepted from the video synthesis unit 151 that is generating a synthesized video including the content video V2. The extracted scene of the content at the time when the input of the feedback information was accepted may be acquired from the playback processing unit 40 that is executing the playback processing of the content.
[0088] In step S3, the acquisition / recording / output control unit 153 associates the acquired content-related information with the feedback information from the input receiving units 152a, 152b, and 152c, and records the information in the feedback DB 160.
[0089] Then, in step S4, the acquisition / recording / output control unit 153 outputs the feedback information associated with the content-related information to the reference UI / UX 200. The feedback information may be output to the reference UI / UX 200 in real time to the user who is reviewing the content, or may be output to the creator or editor of the content after the reviewing process is completed.
[0090] (Example of Input UI) An example of an input UI for feedback information realized by the input receiving units 152a, 152b, and 152c will be described.
[0091] (a) UI for Inputting Selection Information The UI for inputting selection information representing options selected by the user, which is realized by the input receiving unit 152a, will be described with reference to FIG.
[0092] When inputting selection information for a particular scene in the content, the user takes a screenshot of a composite image (CG image) that includes the scene.
[0093] 6, an input button b11 for opening an input screen for inputting feedback information is displayed on the scene image SS1 captured as a screenshot. The input button b11 may be included as part of another display function such as a separately displayed menu panel, in addition to being displayed on the scene image SS1.
[0094] When the input button b11 is selected, a menu screen c12 for selecting a category of content to which an evaluation comment is to be assigned in the review process is displayed on the scene image SS1, as shown in the lower left of Fig. 6. "VFX," "Color," "Cut," etc. are displayed as editing categories for content video. "Dialog," "FX / SE / Foley," "Music," etc. are displayed as editing categories for content audio. The list of categories displayed on the menu screen c12 may be embedded in the application or may be replaceable by a setting file containing template data.
[0095] When one of the categories of content to be evaluated is selected, a menu screen c13 for selecting the category of evaluation comments for the content and the content of the evaluation comments is displayed on the scene image SS1, as shown in the upper right corner of FIG. 6 . The evaluation comment categories displayed include "Positive," "Negative," "Amendment," "Suggestion," and "Question." These categories may be used for sorting on the reference UI / UX 200 so that the editor (creator) can prioritize revision work. The list of categories displayed on the menu screen c13 may be embedded in the application or may be replaceable by a configuration file containing template data.
[0096] Furthermore, the menu screen c13 displays adjectives that represent the content of the selected evaluation comment according to the category of the evaluation comment. Here, candidate adjectives are suggested using a predictive conversion dictionary entered via an input device such as a keyboard. In the example of FIG. 6, "Great" and "Gorgeous" are searched and suggested as adjectives beginning with "G." A general dictionary prepared in advance may be used as the predictive conversion dictionary, or a user dictionary customized for each user may be used preferentially.
[0097] When input (selection) of all evaluation comments is completed, a confirmation screen r14 for confirming the input contents is displayed on the scene image SS1, as shown in the lower right of Fig. 6. In the example of Fig. 6, the time information of the scene when the comment input was accepted, "0:45", and the recorded comment (selection information) "Sound (FX / SE / Foley) / Positive / Great" are displayed as the input contents on the confirmation screen r14. The display time of the confirmation screen r14 may be changed by settings, or it may be set to be hidden.
[0098] (b) Input UI for Selected Information The input UI for entered information based on user input, which is realized by the input receiving unit 152b, will be described with reference to FIG.
[0099] When inputting information for a specific scene in the content, the user takes a screenshot of a composite image (CG image) that includes the scene.
[0100] As shown in Fig. 7, characters and figures can be directly written as written information in a scene image SS2 captured as a screenshot using an input device such as a pen tablet or a mouse. In the example of Fig. 7, the written information includes an arrow f21 indicating a location in the content video that the user wants to point out, and a character string f22 indicating the content of the point indicated by the arrow f21.
[0101] The input device is not limited to a device that allows handwritten input, such as a pen tablet or a mouse, but may also be a device that allows input by selecting characters, such as a keyboard. In this case, a character string consisting of the selected characters may be displayed in the scene image SS2 as written information.
[0102] The storage format of the written information may be image data including characters and figures written directly into the content, such as JPG, PNG, or binary format, or it may be text data extracted from the image data using an OCR (Optical Character Reader).
[0103] The written information may be input into the reference space RS reproduced by the spatial image R3, which forms the background of the content image.
[0104] 8, a group of figures f23 may be written in the reference space RS in a scene image SS2 captured as a screenshot, indicating that the sound source from which the sound effect is output should be moved from the sound source in the middle left of the screen to the sound source in the upper center of the screen. In this case, it becomes possible to easily and explicitly add suggestions or instructions for modifying the sound effects corresponding to the scene, which would be difficult to convey using text alone.
[0105] In addition, if there are parts of the spatial image R3, which is a CG image, where data does not exist, a virtual object representing the size of the reference space RS may be displayed as a complement, such as an icon showing the distance to the walls or ceiling of the reference space RS.
[0106] (c) Input UI for Voice Input Information The input UI for voice input information based on user voice, which is realized by the input receiving unit 152c, will be described with reference to FIG.
[0107] When inputting voice input information for a specific scene in content, the user takes a screenshot of a composite image (CG image) that includes the scene.
[0108] 9, a text box t31 can be displayed in a scene image SS3 captured as a screenshot by a predetermined operation (user input or user voice), and text data converted from the user's speech using voice recognition technology can be entered. In this case, it becomes possible to easily add correction suggestions or correction instructions for the scene without using an input device such as a pen tablet or mouse.
[0109] According to the above-described feedback information input UI, for example, when a user is wearing an HMD to review content, even if the user has difficulty seeing what is in front of them, they can easily record comments, input work instructions, etc. In addition, although the above description assumes that input information is input into a scene image captured as a screenshot of a synthetic image (CG image), input information may also be input to any point in the spatial image (CG image) reproduced by the HMD using an input device such as a pen tablet or mouse, or a VR (Virtual Reality) device such as a controller or haptic glove that operates in conjunction with the HMD.
[0110] 4. Examples of Use of Feedback Information> A description will now be given of examples of use of feedback information received as input from a user viewing content. The feedback information is associated with time information when the input was received and a capture of the content, and is recorded in the feedback DB 160. This makes it possible to achieve more optimal content review and production work.
[0111] (4-1. Display of Feedback Information) Feedback information about the content to be reviewed, which is recorded in the feedback DB 160, is displayed in the reference UI / UX 200. When feedback information is input by multiple users for each scene of the content to be reviewed, feedback information associated with each of multiple pieces of time information is displayed as a list.
[0112] For example, the list includes, in addition to feedback information associated with time information, input form information indicating the input form of the feedback information, version information of the content at the time the feedback information was received, and user information indicating the user who input the feedback information as records.
[0113] FIG. 10 is a diagram showing an example of a list of feedback information.
[0114] 10, time information is associated with each of the following information: type information, category information, audio information, 3D input information, text information, version information, unread information, responded information, and comment owner information. Here, each of the category information, audio information, 3D input information, and text information corresponds to feedback information.
[0115] The time information is the time information when the input of the feedback information is accepted.
[0116] The type is input form information that indicates the input form of the feedback information. "a" indicates that the feedback information was input by selecting a category. "b" indicates that the feedback information was input by writing. "c" indicates that the feedback information was input by voice input.
[0117] The category information may be selection information when feedback information is input by selecting a category, or may be classification information indicating a category determined based on text information, which will be described later.
[0118] The voice information is information indicating whether or not the feedback information has been input by voice input.
[0119] The 3D input information is information indicating whether or not feedback information has been input by handwriting into the CG image portion that forms the background of the content image V2.
[0120] The text information is selected information, written information, or voice input information, which corresponds to the case where the feedback information is input by selection, writing, or voice input.
[0121] The version information is the version information of the content at the time the feedback information is received. A version is assigned to video and audio content at each stage of production and editing, and the quality of the content improves as the version is changed.
[0122] The unread information is information indicating whether or not the feedback information input for the scene corresponding to the time information has been checked by the producer or editor.
[0123] The response completion information is information indicating whether or not the matters pointed out in the feedback information input for the scene corresponding to the time information have been addressed by the producer or editor.
[0124] The comment owner information is user information that indicates a user who inputs feedback information for the scene corresponding to the time information.
[0125] In the list of Figure 10, at time 1:00 (1 minute 00 seconds) for the editor's cut content, Mr. A entered feedback information (selection information) about the dialogue in the content audio by selecting a category, and it is shown that the issues pointed out have been addressed by the producer and editor.
[0126] Furthermore, at 2:30 (2 minutes 30 seconds) in the rough mix content, Mr. B wrote down feedback information (written information) about the content audio and entered it into the CG image, indicating that the points he pointed out were confirmed by the producers and editors.
[0127] Similarly, at 5:00 (5 minutes 00 seconds) in the rough mix content, Mr. C inputs feedback information (voice input information) about the brightness of the content image via voice input, indicating that the points raised have been confirmed by the producers and editors.
[0128] In the list shown in Fig. 10, only desired information can be displayed by filtering or sorting using the filter button b101. This allows comments on a specific version or comments entered by a specific user to be displayed. It is also possible to display comments on a specific category (video / audio) or to display only comments that have not yet been addressed by the producer or editor.
[0129] Here, the UI processing in the reference UI / UX 200 where the above-mentioned list is displayed will be described with reference to the flowchart in Fig. 11. The processing in Fig. 11 is executed under the control of the acquisition / recording / output control unit 153, for example.
[0130] In step S11, the acquisition / recording / output control unit 153 displays a list such as that described with reference to FIG. 10 on the reference UI / UX 200 based on the feedback information recorded in the feedback DB 160.
[0131] In step S12, the acquisition / recording / output control unit 153 determines whether any record has been designated (selected) in the list displayed on the reference UI / UX 200. The record here includes one piece of feedback information and each piece of information corresponding to it.
[0132] Step S12 is repeated until one of the records is designated, and once one of the records is designated, the process proceeds to step S13.
[0133] In step S13, the acquisition / recording / output control unit 153 plays back the content around the time when the comment (feedback information) included in the record specified in the list was added.
[0134] FIG. 12 shows an example of content playback corresponding to a specified record.
[0135] Fig. 12 shows the reference UI / UX 200 when feedback information input for the rough mix content at time 2:30 (2 minutes 30 seconds) is specified in the list of Fig. 10. The screen displayed on the reference UI / UX 200 has a content playback area R111 and a feedback information display area R112.
[0136] The content playback area R111 is an area in which content (content video, content audio) around the time when a comment (feedback information) included in a record specified in the list was added is played. The example of Fig. 12 shows how content is played back for 10 seconds before and after 2:30, the time when the comment was added. The playback time of the content in the content playback area R111 is not limited to the 10 seconds before and after the time when the comment was added, and can be changed as needed.
[0137] The feedback information display area R112 is an area where a comment (feedback information) included in a record specified in the list is displayed. In the example of Fig. 12, the comment "The volume is..." is displayed as text information shown in the list of Fig. 10, and the fact that the comment was input to the CG image is displayed.
[0138] As described above, scenes with comments specified in the feedback information list are played back, allowing producers and editors to quickly check and respond to scenes that require correction.
[0139] (4-2. Changing the Version of Content) As described above, the audio-video content to be reviewed in the review system 1 is assigned a version for each production and editing stage, and the level of completion increases through changes in the version. Changes to the version of content can include, for example, changes to the version of video cuts leading up to post-production, such as editor's cut, director's cut, and final cut, as well as replacing audio editing data and changing the TempMix.
[0140] When the version of a content is changed, the time information of the video section of the content before the version change is also changed. In this case, it is necessary to update the time information associated with the feedback information input for a scene of the content before the version change so that the comment added to that scene can be confirmed.
[0141] Here, the process of updating the time information associated with the feedback information will be described with reference to the flowchart in Fig. 13. The process in Fig. 13 is executed under conditions in which the acquisition / recording / output control unit 153 can access and read the input source file, which is the content after the version has been changed. At this time, the input source file may be recorded locally in the review system 1, or may be downloaded or streamed from a server on the cloud, etc.
[0142] In step S21, the acquisition / recording / output control unit 153 acquires the content whose version has been changed.
[0143] In step S22, the acquisition / recording / output control unit 153 reads out the capture (extracted scene) acquired before the version was changed, based on the feedback information recorded in the feedback DB 160.
[0144] In step S23, the acquisition / recording / output control unit 153 identifies a scene in the content whose version has been changed that corresponds to the read capture (extracted scene). Here, for example, a frame identical to the read capture may be searched for among the frames of the content whose version has been changed. Furthermore, in consideration of the case where the scene in the content whose version has been changed has been edited, a frame that is highly similar to the read capture may be extracted.
[0145] In step S24, the acquisition / recording / output control unit 153 updates the time information of the read capture to the time information of the scene identified in the content whose version has been changed. That is, the time information associated with the feedback information input for the scene in the content before the version was changed is updated.
[0146] This makes it easy to check whether concerns and comments such as correction suggestions and correction instructions that were given before the version of the content was changed have been resolved or addressed.
[0147] Then, in step S25, the acquisition / recording / output control unit 153 assigns a new file name to the session file containing the identified scene and records it. Content is recorded in session files consisting of a certain time unit. By assigning a unique ID as a file name that can identify the session file containing the scene to which feedback information has been assigned and its version, it becomes possible to retroactively refer to feedback information for scenes in past versions.
[0148] (4-3. Application to Computer Games) The content to be reviewed in the review system 1 may be not only the above-mentioned audiovisual content but also a computer game.
[0149] Here, the processing for each game scene in a computer game will be described with reference to the flowchart in Fig. 14. The processing in Fig. 14 is executed under conditions where streaming is possible through two-way communication between a development environment (production environment) such as a computer game console or PC and a review environment where review work is performed.
[0150] In step S31, the input receiving units 152a, 152b, and 152c receive input of feedback information for the computer game.
[0151] In step S32, the acquisition / recording / output control unit 153 acquires, as content-related information, asset data and its configuration information for the game scene for which feedback information has been received. The asset data here refers to various types of material data required for creating a computer game, such as 3DCG model data, animation data, and sound data.
[0152] In step S33, the acquisition / recording / output control unit 153 associates the acquired asset data and its configuration information with the input feedback information, and transmits and records them to the production environment side.
[0153] At this time, the asset data associated with the feedback information and its configuration information are added to the feedback information list as described with reference to Fig. 10. That is, information necessary for reproducing the scene in question during debugging, such as which stage in the game engine and which 3DCG model the feedback information is for, and under what conditions the feedback information is valid, is added.
[0154] This allows the developer (creator) to use this information to reconstruct the scene based on the configuration information on the game engine and reproduce the game scene with feedback information added.
[0155] At this time, graphic information drawn on the screen and input as feedback information, or text information input as a comment may be embedded in the configuration information on the game engine.
[0156] (4-4. Creating a User Dictionary) Feedback information from a specific user (hereinafter referred to as a specific user) may be collected to create a user dictionary that is optimized for individual users.
[0157] The user dictionary creation process will be described with reference to the flowchart of FIG.
[0158] In step S41 , the acquisition / recording / output control unit 153 collects feedback information of a specific user from the feedback information recorded in the feedback DB 160 .
[0159] In step S42, the acquisition / recording / output control unit 153 creates a user dictionary that is personalized for the specific user based on the collected feedback information.
[0160] For example, as shown in FIG. 16, by using a list of feedback information, the comment owner can collect only the feedback information of person A, thereby creating a user dictionary UD that is personally optimized for person A.
[0161] The collected feedback information includes expressions such as words and adjectives that are often used as comments on content. In other words, the user dictionary UD can register words and expressions that Mr. A often uses, and these can be presented as fixed phrase data that are options related to content.
[0162] Returning to the flowchart of FIG. 15, in step S43, the acquisition / recording / output control unit 153 performs machine learning on combinations of collected feedback information and scenes of content to which each piece of feedback information has been input.
[0163] Based on the scenes of the content that have been machine-learned in this way, candidate information that is a candidate for feedback information to be input by a specific user may be presented.
[0164] Therefore, the candidate information presentation process will be described with reference to the flowchart in Fig. 17. The process in Fig. 17 is executed, for example, in a situation where a specific user who has created a user dictionary is performing a review operation on new content as a review target.
[0165] In step S51, the acquisition / recording / output control unit 153 determines whether a scene similar to a machine-learned scene (i.e., a scene in the content when the input of the feedback information used to create the user dictionary was accepted) has been input as a scene in the content to be reviewed.
[0166] Step S51 is repeated until a scene similar to the machine-learned scene is input, and once a scene similar to the machine-learned scene is input, the process proceeds to step S52.
[0167] In step S52, the acquisition / recording / output control unit 153 automatically generates candidate information based on the machine-learned scenes. Specifically, the acquisition / recording / output control unit 153 automatically generates candidate information using feedback information combined with the machine-learned scenes. The automatically generated candidate information may be a list of adjectives or sentences that are expected to be input by a specific user.
[0168] In step S53 , the acquisition / recording / output control unit 153 presents the automatically generated candidate information to the specific user via the video synthesis unit 151 , the display controller 50 , and the display 60 .
[0169] As described above, by creating a personally optimized user dictionary and presenting candidate information based on feedback information used to create the user dictionary, it is possible to simplify review work, such as allowing the user to add comments.
[0170] 5. Example of Computer Hardware Configuration The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware, a general-purpose personal computer, or the like.
[0171] 18 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes using a program. The information processing device 100 is, for example, configured by a computer 500 having a configuration similar to that shown in FIG.
[0172] A CPU 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .
[0173] An input / output interface 505 is also connected to the bus 504. An input unit 506 including a keyboard, a mouse, etc., and an output unit 507 including a display, a speaker, etc. are connected to the input / output interface 505. Also connected to the input / output interface 505 are a storage unit 508 including a hard disk, a nonvolatile memory, etc., a communication unit 509 including a network interface, etc., and a drive 510 that drives removable media 511.
[0174] In the computer 500 configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program stored in the memory unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.
[0175] The program executed by the CPU 501 is installed in the storage unit 508 by being recorded on, for example, a removable medium 511 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.
[0176] The program executed by computer 500 may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0177] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0178] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0179] The embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.
[0180] For example, the embodiment of the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.
[0181] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0182] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0183] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0184] Furthermore, the technology according to the present disclosure may have the following configurations: (1) An information processing system including: an input accepting unit that accepts input information related to content from a user viewing the content; an acquiring unit that acquires content-related information when the input information is accepted; a recording control unit that records the input information in association with the content-related information; and an output control unit that outputs the input information associated with the content-related information. (2) The information processing system described in (1), in which the content-related information includes at least one of time information of the content when the input information is accepted and an extracted scene of the content. (3) The information processing system described in (1) or (2), in which the input information includes at least one of selection information related to the content, written information for the content, and voice input information about the content. (4) The information processing system described in (3), in which the selection information includes at least one of a category of the content selected by the user, a category of evaluation for the content, and details of the evaluation. (5) The information processing system according to (3), wherein the entry information includes at least one of image data including characters and figures entered directly in the content and text data extracted from the image data. (6) The information processing system according to (3), wherein the voice input information includes text data converted from the user's speech. (7) The information processing system according to any one of (1) to (6), wherein the content includes content audio played back using sound transfer characteristics measured in a reference space, and a composite video obtained by combining a content video corresponding to the content audio with a spatial video based on the reference space. (8) The information processing system according to (7), wherein the input accepting unit accepts input of the input information by the user wearing an HMD (Head Mounted Display) and viewing the composite video. (9) The information processing system according to (1), wherein the output control unit displays the input information associated with each of the plurality of content-related information as a list.(10) The information processing system described in (9), wherein the list includes, as records, input form information indicating the input form of the input information, version information of the content when the input information was accepted, and user information indicating the user who entered the input information, in addition to the input information associated with the content-related information. (11) The information processing system described in (10), wherein, when one of the records is designated in the list, the output control unit plays the content around the time when the input information included in the designated record was accepted. (12) The information processing system described in (2), wherein, when the version of the content is changed, the output control unit updates the time information associated with the input information. (13) The information processing system described in (12), wherein the output control unit updates the time information by identifying a scene in the content whose version has been changed that corresponds to the extracted scene acquired before the version was changed. (14) The information processing system described in (1), wherein the recording control unit records the input information associated with the content-related information in a content production environment. (15) The information processing system according to (14), wherein the content is a computer game, and the content-related information includes asset data of game scenes of the computer game and configuration information thereof. (16) The information processing system according to (1), wherein the output control unit creates a user dictionary that is personalized for the specific user based on the input information input by the specific user. (17) The information processing system according to (16), wherein the output control unit presents candidate information that is a candidate for the input information to be input by the specific user based on the input information used to create the user dictionary. (18) The information processing system according to (17), wherein the output control unit automatically generates the candidate information based on a scene of the content at the time when the input information used to create the user dictionary was accepted.(19) An information processing device comprising: an output control unit that, from among input information about content received from a user viewing the content and content-related information at the time the input of the input information was received, outputs the input information that is associated with the content-related information, among input information about the content received from the user viewing the content and content-related information at the time the input of the input information was received, in association with each other and recorded. (20) A program that causes a computer to execute a process of: outputting the input information that is associated with the content-related information, among input information about the content received from the user viewing the content and content-related information at the time the input of the input information was received, in association with each other and recorded.
[0185] 1 Review system, 100 Information processing device, 150 Review application, 151 Video synthesis unit, 152a, 152b, 152c Input reception unit, 153 Acquisition / recording / output control unit, 160 Feedback DB, 200 Reference UI / UX
Claims
1. An information processing system including: an input receiving unit that receives input information related to content from a user viewing the content; an acquisition unit that acquires content-related information when the input information is received; a recording control unit that records the input information in association with the content-related information; and an output control unit that outputs the input information associated with the content-related information.
2. The information processing system according to claim 1, wherein the content-related information includes at least one of time information of the content when the input information was received and an extracted scene of the content.
3. The information processing system according to claim 1, wherein the input information includes at least one of selection information regarding the content, written information for the content, and voice input information for the content.
4. The information processing system according to claim 3, wherein the selection information includes at least one of the category of the content selected by the user, the category of a rating for the content, and the details of the rating.
5. The information processing system according to claim 3, wherein the written information includes at least one of image data including characters and figures written directly into the content, and text data extracted from the image data.
6. The information processing system according to claim 3, wherein the voice input information includes text data converted from the user's spoken voice.
7. The information processing system of claim 1, wherein the content includes content audio reproduced using sound transmission characteristics measured in a reference space, and a composite image obtained by combining a content image corresponding to the content audio with a spatial image based on the reference space.
8. The information processing system according to claim 7, wherein the input receiving unit receives the input of the input information from the user wearing an HMD (Head Mounted Display) and viewing the composite video.
9. The information processing system according to claim 1, wherein the output control unit displays the input information associated with each of the plurality of pieces of content-related information as a list.
10. An information processing system as described in claim 9, wherein the list includes, as records, in addition to the input information associated with the content-related information, input form information indicating the input form of the input information, version information of the content at the time the input information was accepted, and user information indicating the user who entered the input information.
11. The information processing system according to claim 10, wherein when one of the records in the list is specified, the output control unit plays the content near the time when the input information included in the specified record was received.
12. The information processing system according to claim 2, wherein the output control unit updates the time information associated with the input information when the version of the content is changed.
13. The information processing system of claim 12, wherein the output control unit updates the time information by identifying a scene in the content whose version has been changed that corresponds to the extracted scene obtained before the version was changed.
14. The information processing system according to claim 1, wherein the recording control unit records the input information associated with the content-related information in the content production environment.
15. The information processing system according to claim 14, wherein the content is a computer game, and the content-related information includes asset data of game scenes of the computer game and configuration information thereof.
16. The information processing system according to claim 1, wherein the output control unit creates a user dictionary that is personalized for the specific user based on the input information entered by the specific user.
17. The information processing system according to claim 16, wherein the output control unit presents candidate information that is a candidate for the input information to be input by the specific user based on the input information used to create the user dictionary.
18. The information processing system according to claim 17, wherein the output control unit automatically generates the candidate information based on a scene of the content when the input information used to create the user dictionary was received.
19. An information processing device comprising an output control unit that outputs input information associated with content-related information from among input information about the content received from a user viewing the content and content-related information at the time the input information was received, the input information being associated with the content-related information and stored in association with the content.
20. A program for causing a computer to execute a process of outputting input information associated with content-related information from among input information related to content received from a user viewing the content and content-related information recorded in association with the input of the input information.
Citation Information
Patent Citations
Content evaluation system
JP2007334704A
Moving image reproduction apparatus, candidate extraction method, and program
JP2015211291A
Acoustic signal processing device and program
JP2023122230A
Information processing device, information processing method, and information processing program
WO2020079996A1