Information processing program, information processing method, and information processing system
The system addresses the lack of personalized content delivery by using geofence areas and sensor data to adjust activation conditions, improving user experience through location-specific content presentation.
Patent Information
- Application Number
- JP2024085476
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-20
- Filing Date
- 2024-05-27
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2040-06-08
AI Technical Summary
Existing information processing systems fail to provide a better user experience when delivering context-based services, as they do not effectively utilize geofence areas and sensor data to tailor content element activation conditions for individual users.
An information processing system that includes an interface for displaying geofence areas and receiving inputs to change activation conditions, along with a control unit that adjusts these conditions based on user location and sensor data, allowing for personalized content element presentation.
Enables a better user experience by delivering context-aware content elements tailored to individual user locations, enhancing the engagement and relevance of services.
Smart Images

Figure 0007732538000001 
Figure 0007732538000002 
Figure 0007732538000003
Abstract
Description
[Technical Field]
[0001] This technology is Information processing program, information processing method, and Regarding information processing systems, in particular, it has been made possible to provide a better user experience. Information processing program, information processing method, and Regarding information processing systems. [Background technology]
[0002] BACKGROUND ART In recent years, with the widespread use of information devices, various services that take advantage of the characteristics of the devices have been provided (see, for example, Patent Document 1).
[0003] In this type of service, processing may be performed using context information. Known technologies relating to context include those disclosed in Patent Documents 2 to 5. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 6463529 [Patent Document 2] Japanese Patent Application Laid-Open No. 2015-210818 [Patent Document 3] International Publication No. 2013 / 136792 [Patent Document 4] Japanese Patent Application Laid-Open No. 2007-172524 [Patent Document 5] International Publication No. 2016 / 136104 Summary of the Invention [Problem to be solved by the invention]
[0005] However, when providing services using context information, it is necessary to provide a better user experience.
[0006] The present technology was developed in light of these circumstances, and aims to provide a better user experience. [Means for solving the problem]
[0007] One aspect of this technology Information Processing Program teeth, An information processing program for causing a computer to function as an information processing device comprising: an interface unit that displays a geofence area corresponding to a first activation condition that is a condition for playing a content element and receives input; and a control unit that changes the first activation condition based on an input of a change to the geofence area and stores, in a storage unit, an activation condition including the changed first activation condition for the content element based on the change to the first activation condition. is.
[0008] One aspect of this technology Information Processing Program teeth, an information processing program for causing a computer to function as an information processing device, the information processing device comprising: an acquisition unit that acquires sensor data, in which context information is associated with content elements and activation conditions are associated with the context information; and a control unit that, when the sensor data satisfies an activation condition that is a condition for playing a content element, controls to output the content element associated with the context information corresponding to the activation condition, the sensor data including a position of a user or a device used by a user, the activation condition including a condition related to an activation range within which the output of the content element is controlled, and the control unit controls the output of the content element corresponding to the activation condition according to the position of the user or the device used by a user and the activation condition, and controls a sound image localization position of a sound source within the content element. is.
[0009] One aspect of this technology Information Processing Method teeth, An information processing method including: an information processing device displaying a geofence area corresponding to a first activation condition, which is a condition for playing a content element, and receiving an input; changing the first activation condition based on the input of a change to the geofence area; and storing, in a storage unit, an activation condition for the content element, including the changed first activation condition, based on the change to the first activation condition. is.
[0010] One aspect of this technology Information Processing Systems teeth, An information processing system comprising: an interface unit that displays a geofence area corresponding to a first activation condition, which is a condition for playing a content element, and receives input; and a control unit that changes the first activation condition based on an input of a change to the geofence area, and stores an activation condition including the changed first activation condition for the content element in a storage unit based on the change to the first activation condition. do. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a representative diagram illustrating an overview of the present technology. [Figure 2] FIG. 1 is a diagram illustrating an example of the configuration of an information processing system to which the present technology is applied. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of a data management server in FIG. 2. [Figure 4] FIG. 3 is a diagram illustrating an example of the configuration of the editing device in FIG. 2. [Figure 5] FIG. 3 is a diagram illustrating an example of the configuration of the playback device in FIG. 2. [Figure 6] FIG. 1 is a diagram illustrating an overall image of information processing in a first embodiment. [Figure 7] 4 is a flowchart illustrating a detailed flow of information processing in the first embodiment. [Figure 8] FIG. 10 is a diagram illustrating an example of information stored in a scenario DB. [Figure 9] FIG. 10 is a diagram illustrating an example of information stored in a user scenario DB. [Figure 10] FIG. 10 is a diagram illustrating another example of information stored in the scenario DB. [Figure 11] FIG. 2 is a diagram illustrating an example of a content element. [Figure 12]FIG. 10 is a diagram illustrating an example of a combination of a content element and a context. [Figure 13] FIG. 10 is a diagram illustrating an example of a scenario. [Figure 14] FIG. 10 is a diagram showing an example of a scenario selection / new creation screen. [Figure 15] FIG. 10 is a diagram illustrating an example of a scenario editing screen. [Figure 16] FIG. 10 is a diagram showing a first example of a geofence editing screen. [Figure 17] FIG. 10 is a diagram showing a second example of a geofence editing screen. [Figure 18] FIG. 10 is a diagram illustrating an overall image of information processing in a second embodiment. [Figure 19] FIG. 11 is a diagram showing an overall image of information processing in a third embodiment. [Figure 20] FIG. 10 is a diagram showing an example of setting an activation condition for content element-context information. [Figure 21] FIG. 10 is a diagram showing an example of a scenario selection / playback screen. [Figure 22] FIG. 10 is a diagram showing an example of an activation condition setting screen. [Figure 23] FIG. 10 is a diagram showing an example of an activation condition detailed setting screen. [Figure 24] FIG. 10 is a diagram showing an example of a content element selection screen. [Figure 25] FIG. 10 is a diagram showing an example of a content element editing screen. [Figure 26] FIG. 10 is a diagram illustrating an example of a scenario selection screen. [Figure 27] FIG. 10 is a diagram showing a first example of an activation condition setting screen. [Figure 28] FIG. 10 is a diagram showing a second example of an activation condition setting screen. [Figure 29] FIG. 10 is a diagram illustrating an example of a geofence editing screen. [Figure 30] FIG. 10 is a diagram illustrating an example of user scenario settings. [Figure 31] FIG. 13 is a diagram showing an overall image of information processing in a fourth embodiment. [Figure 32]FIG. 13 is a diagram showing an overall image of information processing in a fourth embodiment. [Figure 33] FIG. 10 is a diagram showing an example of a combination of activation conditions and sensing means. [Figure 34] FIG. 10 is a diagram illustrating an example of a state in which activation conditions overlap. [Figure 35] FIG. 10 is a diagram showing a first example of how to respond when activation conditions overlap. [Figure 36] FIG. 10 is a diagram showing a second example of how to respond when activation conditions overlap. [Figure 37] FIG. 10 is a diagram showing a third example of how to respond when activation conditions overlap. [Figure 38] FIG. 10 is a diagram showing a fourth example of how to respond when activation conditions overlap. [Figure 39] FIG. 10 is a diagram illustrating an example of the configuration of an information processing system when multiple characters are arranged. [Figure 40] FIG. 10 is a diagram showing an example of information stored in a character placement DB. [Figure 41] FIG. 10 is a diagram illustrating an example of information stored in a location-dependent information DB. [Figure 42] FIG. 10 is a diagram illustrating an example of information stored in a scenario DB. [Figure 43] FIG. 10 is a diagram showing a first example of a multiple character arrangement. [Figure 44] FIG. 10 is a diagram showing a second example of a multiple character arrangement. [Figure 45] FIG. 10 is a diagram showing a third example of a multiple character arrangement. [Figure 46] FIG. 20 is a diagram showing an overall image of information processing in the sixth embodiment. [Figure 47] FIG. 20 is a diagram showing an overall picture of information processing in the seventh embodiment. [Figure 48] FIG. 20 is a diagram showing an overall picture of information processing in the eighth embodiment. [Figure 49] FIG. 13 is a diagram showing an overall picture of information processing in the ninth embodiment. [Figure 50] FIG. 23 is a diagram showing an overall picture of information processing in the tenth embodiment. [Figure 51] FIG. 23 is a diagram showing an overall picture of information processing in the eleventh embodiment. [Figure 52] FIG. 1 illustrates an example of the configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present technology will be described with reference to the drawings. The description will be made in the following order.
[0013] 1. First embodiment: Basic configuration 2. Second embodiment: Creation of a scenario DB 3. Third embodiment: Generation of different media 4. Fourth embodiment: Creation of user scenario DB 5. Fifth embodiment: Configuration of sensing means 6. Sixth embodiment: Configuration when activation conditions are set for multiple context information 7. Seventh embodiment: Configuration in which multiple devices are linked 8. Eighth embodiment: Configuration in cooperation with other services 9. Ninth embodiment: Scenario sharing configuration 10. Tenth embodiment: Other examples of data 11. Eleventh embodiment: Configuration using user feedback 12. Variations 13. Computer Configuration
[0014] (Representative image) FIG. 1 is a representative diagram showing an overview of this technology.
[0015] This technology allows users living in different locations to access the same scenario, providing a better user experience.
[0016] In Figure 1, a creator creates a scenario by adding context information to content elements, which are the elements that make up the content, using editing equipment such as a personal computer. The scenario created in this way is distributed via a server on the Internet.
[0017] Each user creates their own user scenario by operating a playback device such as a smartphone to select the desired scenario from the distributed scenarios and set the activation conditions, which are the conditions for presenting content elements. In other words, in Figure 1, two users, User A and User B, each set their own activation conditions for the same scenario, so the activation conditions for the user scenario differ for each user.
[0018] Therefore, the same scenario is executed for each user in different locations, and one scenario can be used by users living in different locations.
[0019] <1. First embodiment>
[0020] (System configuration example) FIG. 2 shows an example of the configuration of an information processing system to which this technology is applied.
[0021] The information processing system 1 is made up of a data management server 10, an editing device 20, and playback devices 30-1 to 30-N (N: an integer equal to or greater than 1). In the information processing system 1, the data management server 10, the editing device 20, and the playback devices 30-1 to 20-N are interconnected via the Internet 40.
[0022] The data management server 10 is composed of one or more servers for managing data such as a database, and is installed in a data center or the like.
[0023] The editing device 20 is configured from information devices such as personal computers and is managed by a service provider. The editing device 20 connects to the data management server 10 via the Internet 40, performs editing processing on data stored in a database, and generates a scenario.
[0024] The playback device 30-1 is composed of information devices such as smartphones, mobile phones, tablet terminals, wearable devices, portable music players, game consoles, and personal computers.
[0025] The playback device 30-1 connects to the data management server 10 via the Internet 40, and generates a user scenario by setting trigger conditions for the scenario. Based on the user scenario, the playback device 30-1 plays back content elements in accordance with the trigger conditions.
[0026] Like the playback device 30-1, the playback devices 30-2 to 30-N are configured as information devices such as smartphones, and play back content elements according to the trigger conditions based on the generated user scenario.
[0027] In the following description, the playback devices 30-1 to 20-N will be simply referred to as playback device 30 unless there is a need to distinguish between them.
[0028] (Data management server configuration example) FIG. 3 shows an example of the configuration of the data management server 10 of FIG.
[0029] 3, the data management server 10 includes a control unit 100, an input unit 101, an output unit 102, a storage unit 103, and a communication unit 104.
[0030] The control unit 100 is configured with a processor such as a CPU (Central Processing Unit), etc. The control unit 100 is a central processing device that controls the operations of each unit and performs various types of arithmetic processing.
[0031] The input unit 101 is configured with a mouse, a keyboard, physical buttons, etc. The input unit 101 supplies the control unit 100 with an operation signal corresponding to a user operation.
[0032] The output unit 102 is configured with a display, a speaker, etc. The output unit 102 outputs video, audio, etc. under the control of the control unit 100.
[0033] The storage unit 103 is configured by a large-capacity storage device such as a semiconductor memory including a nonvolatile memory or a volatile memory, a hard disk drive (HDD), etc. The storage unit 103 stores various types of data under the control of the control unit 100.
[0034] The communication unit 104 is configured with a communication module that supports wireless or wired communication conforming to a predetermined standard, etc. The communication unit 104 communicates with other devices under the control of the control unit 100.
[0035] The control unit 100 also includes a data management unit 111 , a data processing unit 112 , and a communication control unit 113 .
[0036] The data management unit 111 manages various databases and content data stored in the storage unit 103 .
[0037] The data processing unit 112 performs data processing on various types of data, including processing on content and processing on machine learning.
[0038] The communication control unit 113 controls the communication unit 104 to exchange various data with the editing device 20 or the playback device 30 via the Internet 40 .
[0039] The configuration of the data management server 10 shown in FIG. 3 is an example, and some of the components may be removed, or other components such as a dedicated image processing unit may be added.
[0040] (Example of editing equipment configuration) FIG. 4 shows an example of the configuration of the editing device 20 of FIG.
[0041] 4, the editing device 20 includes a control unit 200, an input unit 201, an output unit 202, a storage unit 203, and a communication unit 204.
[0042] The control unit 200 is configured with a processor such as a CPU, etc. The control unit 200 is a central processing unit that controls the operation of each unit and performs various types of arithmetic processing.
[0043] The input unit 201 is configured with input devices such as a mouse 221 and a keyboard 222. The input unit 201 supplies the control unit 200 with an operation signal corresponding to a user operation.
[0044] The output unit 202 is configured from output devices such as a display 231 and a speaker 232. The output unit 202 outputs information corresponding to various data under the control of the control unit 200.
[0045] The display 231 displays an image according to the image data from the control unit 200. The speaker 232 outputs an audio (sound) according to the audio data from the control unit 200.
[0046] The storage unit 203 is configured with a semiconductor memory such as a nonvolatile memory, etc. The storage unit 203 stores various data under the control of the control unit 200.
[0047] The communication unit 204 is configured with a communication module that supports wireless or wired communication conforming to a predetermined standard, etc. The communication unit 204 communicates with other devices under the control of the control unit 200.
[0048] The control unit 200 also includes an edit processing unit 211 , a presentation control unit 212 , and a communication control unit 213 .
[0049] The editing processing unit 211 performs editing processing on various types of data, including processing on a scenario, which will be described later.
[0050] The presentation control unit 212 controls the output unit 202 to control the presentation of information such as video and audio corresponding to data such as video data and audio data.
[0051] The communication control unit 213 controls the communication unit 204 to exchange various data with the data management server 10 via the Internet 40 .
[0052] The configuration of the editing device 20 shown in FIG. 4 is an example, and some of the components may be removed or other components may be added.
[0053] (Example of playback device configuration) FIG. 5 shows an example of the configuration of the playback device 30 of FIG.
[0054] 5, the playback device 30 includes a control unit 300, an input unit 301, an output unit 302, a storage unit 303, a communication unit 304, a sensor unit 305, a camera unit 306, an output terminal 307, and a power supply unit 308.
[0055] The control unit 300 is configured with a processor such as a CPU, etc. The control unit 300 is a central processing unit that controls the operation of each unit and performs various types of arithmetic processing.
[0056] The input unit 301 is configured with input devices such as physical buttons 321, a touch panel 322, a microphone, etc. The input unit 301 supplies the control unit 300 with an operation signal corresponding to a user operation.
[0057] The output unit 302 is configured from output devices such as a display 331 and a speaker 332. The output unit 302 outputs information corresponding to various data under the control of the control unit 300.
[0058] The display 331 displays an image according to the image data from the control unit 300. The speaker 332 outputs an audio (sound) according to the audio data from the control unit 300.
[0059] The storage unit 303 is configured with a semiconductor memory such as a nonvolatile memory, etc. The storage unit 303 stores various data under the control of the control unit 300.
[0060] The communication unit 304 is configured as a communication module that supports wireless communication such as a wireless LAN (Local Area Network), cellular communication (e.g., LTE-Advanced or 5G), or Bluetooth (registered trademark), or wired communication. The communication unit 304 communicates with other devices under the control of the control unit 300.
[0061] The sensor unit 305 is configured from various sensor devices, etc. The sensor unit 305 senses the user and the surroundings, etc., and supplies the control unit 300 with sensor data according to the sensing results.
[0062] Here, the sensor unit 305 can include an inertial sensor that measures position, orientation, acceleration, and speed, a biosensor that measures information such as the heart rate, body temperature, or posture of a living organism, a magnetic sensor that measures the magnitude and direction of a magnetic field, a proximity sensor that measures nearby objects, etc. Note that instead of the inertial sensor, an acceleration sensor that measures acceleration, or a gyro sensor that measures angle (posture), angular velocity, and angular acceleration may be used.
[0063] The camera unit 306 is configured with an optical system, an image sensor, a signal processing circuit, etc. The camera unit 306 supplies the control unit 300 with imaging data obtained by imaging a subject.
[0064] The output terminal 307 is connected via a cable to a device including an electroacoustic conversion device such as an earphone or a headphone. The output terminal 307 outputs data such as audio data from the control unit 300. Note that the connection to a device such as an earphone is not limited to a wired connection, and may be made via wireless communication such as Bluetooth (registered trademark).
[0065] The power supply unit 308 is composed of a battery such as a secondary battery and a power supply management circuit, and supplies power to each unit including the control unit 300 .
[0066] The control unit 300 also includes a playback processing unit 311 , a presentation control unit 312 , and a communication control unit 313 .
[0067] The playback processing unit 311 performs playback processing for various content data, including playback of (part of) music, character speech, and other data.
[0068] The presentation control unit 312 controls the output unit 302 to control the presentation of information such as video and audio corresponding to data such as video data and audio data. The presentation control unit 312 also controls the presentation of data played back by the playback processing unit 311.
[0069] The communication control unit 313 controls the communication unit 304 to exchange various data with the data management server 10 via the Internet 40 .
[0070] Note that the configuration of the playback device 30 shown in FIG. 5 is just an example, and some of the components such as the camera unit 306 and the output terminal 307 may be removed, or other components such as an input terminal may be added.
[0071] The information processing system 1 is configured as described above. The specific details of the information processing executed by the information processing system 1 will be described below.
[0072] (Overall processing) First, with reference to FIG. 6, an overview of information processing in the first embodiment will be described.
[0073] In the data management server 10, the storage unit 103 stores the following databases: a content element-context information DB 151, a scenario DB 152, and a user scenario DB 153. The storage unit 103 also stores data of content elements.
[0074] The content element-context information DB 151 is a database that stores a table that associates content elements with context information.
[0075] Here, content elements are elements that make up content. For example, content elements include dialogue, background music, sound effects, environmental sounds, music, images, and the like that are generated from content such as video and music.
[0076] Context information is information about a context that is assigned to a content element. For example, context information assigned according to a situation in which a content element is expected to be used is associated with the content element and stored in the content element-context information DB 151. Note that, here, the context information may be automatically assigned to the content element using machine learning technology.
[0077] The scenario DB 152 is a database that stores scenarios.
[0078] Here, a scenario is a package of a data set consisting of a combination of content elements and context information (hereinafter also referred to as "content elements-context information") based on a certain theme.
[0079] The scenario DB 152 may store device function information relating to the functions of the playback device 30. By using this device function information, it is possible to execute processing according to the functions of one or more playback devices 30.
[0080] The user scenario DB 153 is a database that stores user scenarios.
[0081] Here, a user scenario is a scenario in which a data set consisting of content elements and context information is packaged, and activation conditions are set for the scenario.
[0082] That is, for each user, it is possible to set activation conditions for at least the context information, and it is possible to generate a user scenario that includes a data set of the context information and the activation conditions. In other words, the user scenario can be said to be a user-defined scenario.
[0083] The triggering condition is the condition for presenting the content element associated with the context information, which is a data set, to the user. The triggering condition can be, for example, a spatial condition such as a position or a location, a temporal condition, or a user's behavior.
[0084] In the information processing system 1, the data management server 10 manages the above database, and the editing device 20 and the playback device 30 access the information stored in the database, thereby performing the processing shown in FIG.
[0085] That is, the playback device 30 senses the user in real time (S101), and it is determined whether the sensor data obtained by the sensing satisfies the triggering conditions set in the user scenario (S102).
[0086] Then, when the sensor data satisfies the activation condition ("Yes" in S102), a content element associated with context information corresponding to the activation condition is presented to the user (S103).
[0087] For example, assume a scenario in which a content element that is "character utterance" is associated with context information that is "home," and an activation condition that is "within a radius of 10 meters from the center of the home" is set for the context information. In this case, based on sensor data (location information), when the user comes within 10 meters of the home, the desired character's utterance is output from the playback device 30 carried by the user.
[0088] (Processing flow) Next, a detailed flow of information processing in the first embodiment will be described with reference to the flowchart of FIG.
[0089] Of the processes shown in FIG. 7, the processes of steps S121 to S127 are mainly processes that are performed when the scenario generation tool is executed by the editing device 20 (control unit 200), and the processes of steps S128 to S133 are mainly processes that are performed when the user scenario generation tool is executed by the playback device 30 (control unit 300).
[0090] In other words, the scenario generation tool is operated by a producer who creates a scenario using editing equipment 20, while the user scenario generation tool is operated by a user who owns playback equipment 30, and the operators of each tool are different, or even if they are the same operator, the timing of operation differs.
[0091] In the editing device 20, the scenario generation tool acquires content (S121), and presents candidates for content elements (S122). Then, content elements are extracted from the content in response to an operation by the creator (S123).
[0092] Furthermore, in the editing device 20, candidates for context information are presented by the scenario generation tool (S124). Then, the context information is assigned to the content elements in accordance with the creator's operation (S125). However, in this case, the context information may be assigned automatically using machine learning technology, not limited to the creator's operation.
[0093] The content elements and context information associated in this way are sent to the data management server 10 and stored in the content element-context information DB 151.
[0094] In the editing device 20, a scenario is generated in accordance with the operation of the creator using the scenario generation tool (S126), and the scenario is saved (S127).
[0095] That is, the scenario generated by the scenario generation tool is sent to the data management server 10 and stored in the scenario DB 152. The scenario stored in the scenario DB 152 can be distributed via the Internet 40.
[0096] Meanwhile, in the playback device 30, the scenario distributed from the data management server 10 is acquired by the user scenario generation tool (S128).
[0097] Then, in the playback device 30, an activation condition is assigned in response to the user's operation (S129). As a result, a user scenario in response to the user's operation is generated from the scenario, and the user scenario is saved (S130).
[0098] The user scenario generated by the user scenario generation tool is sent to the data management server 10 and stored in the user scenario DB 153. This allows the user scenario to be shared with other users.
[0099] Here, if another scenario is to be added ("Yes" in S131), the processes of steps S128 to S130 described above are repeated.
[0100] Furthermore, the playback device 30 can use the user scenario generation tool to launch the created user scenario (S132) and evaluate it (S133).
[0101] The scenario generation tool will be described in detail later with reference to Figures 14 to 17. The user scenario generation tool will be described in detail later with reference to Figures 21 to 25 and Figures 26 to 29.
[0102] The detailed flow of information processing has been explained above.
[0103] (Database example) Next, examples of databases managed by the data management server 10 will be described with reference to FIGS.
[0104] As shown in Fig. 8, the scenario DB 152 stores data sets consisting of combinations of content elements and context information in response to operations of the user scenario generation tool. For example, in Fig. 8, the context information "home" is associated with the content elements "character speech #1" and "BGM #1."
[0105] As shown in FIG. 9, the user scenario DB 153 stores data sets consisting of combinations of content elements and context information, as well as activation conditions assigned to the data sets in response to operations of the user scenario generation tool.
[0106] For example, in Figure 9, the activation conditions of "center (35.631466, 139.743660)" and "radius 10m" are assigned to the content elements "character speech #1" and "BGM #1" and the context information "home." Note that the "a" and "b" in "center (a, b)" refer to the latitude (northern latitude) and longitude (eastern longitude), and represent the activation range of the content element.
[0107] 8 and 9 are merely examples, and other configurations may be used. For example, as shown in Fig. 10, common context information can be assigned to different works (e.g., work A, which is a "movie," work B, which is an "animation," and work C, which is a "literature reading").
[0108] For example, in Figure 10, the context information "Home" is associated with the content elements "BGM #2" of work A, "Character Speech #1" and "BGM #1" of work B, and "Reading #1" of work C.
[0109] The first embodiment has been described above. In this first embodiment, context information is associated with content elements in advance, activation conditions can be set for at least the context information for each user, and a user scenario including a data set of context information and activation conditions can be generated. Then, when sensor data obtained by sensing the user in real time satisfies the activation conditions set in the user scenario, content elements associated with the context information according to the activation conditions are presented to the user.
[0110] This allows each user to enjoy the worldview of the scenario according to the conditions for invoking the user scenario, providing a better user experience.
[0111] <2. Second embodiment>
[0112] The content currently being distributed and streamed includes a variety of formats, such as videos like movies, anime, and games, still images like photographs, paintings, and manga, audio like music and audiobooks, and text like books. However, content that has a story (theatricality) in particular is often made up of elements such as dialogue, effects, and backgrounds.
[0113] When considering the superimposition of the content into the user's everyday space, in addition to presenting the content in the format in which it is distributed or delivered, the content may be re-edited. This re-editing of the content may involve, for example, cutting out a portion of the content in time to fit the spatial and temporal size of the user's current context, or extracting and presenting the above elements to fit the context.
[0114] Below, some of this re-edited content corresponds to the content elements described above. For example, as shown in Figure 11, content elements of a certain piece of content include dialogue, background, music, lyrics, people, symbols, characters, objects, etc.
[0115] The content elements are assigned with the above-mentioned context information, in the form of text, images, audio, etc. Furthermore, the relationship information between the content elements and the context information itself, or a collection of multiple pieces of relationship information, is stored in the scenario DB 152 as a scenario.
[0116] It should be noted that one or more context tags may be assigned to one content element, and the same context tag may be assigned to multiple content elements.
[0117] For example, as shown in Figure 12, from content consisting of video and audio, such as a distributed movie, anime, or game, only the lines of a certain character are extracted to create audio content, and the text "gaining courage" is added as context information to represent the context in which the line is expected to be heard.
[0118] Also, for example, as shown in FIG. 12, a combination of dialogue and background music used in a certain scene is treated as one piece of audio content, and the text "Encounter at the inn" is added as context information.
[0119] Then, the two data sets of "content element-context information" shown in FIG.
[0120] For example, in the case of audio data, dialogue, sound effects, background sounds, background music, etc. are produced as separate multi-track audio sources during production, and then mixed down before distribution and delivery. Therefore, content elements can be extracted from each track before mixing down.
[0121] Furthermore, for example, in the case of images, there is also a technique in which people, backgrounds, objects, etc. are photographed separately and then combined, and content elements can also be extracted from the data before combination.
[0122] The creation of these content elements and the addition of contextual information can be done manually, automatically, or in a combination of the two. The following section focuses on the cases where an automated process is involved.
[0123] Machine learning techniques are available to identify elements such as people, living things, objects, buildings, and landscapes contained in a scene from image information or audio information contained in videos or still images. These techniques can be used to determine the range of content elements and (automatically) generate one or more pieces of context information assumed from the identification results or a combination thereof.
[0124] A data set of "content element-context information" may be automatically generated from this information, or the "content element-context information" may be set manually using this information as reference information.
[0125] A scenario is composed of one or more data sets of "content elements-context information" organized around a certain theme, such as the name of the original work that was re-edited, the characters that appear, the setting, and the emotions that are evoked, and is stored in the scenario DB 152.
[0126] For example, as shown in FIG. 13, the two data sets of "content element-context information" shown in FIG. 12 can be stored in the scenario DB 152 as a scenario of "starting town."
[0127] This allows users to search for and obtain not only the "content element-context information" dataset they want to use, but also multiple "content element-context information" datasets packaged based on a scenario.
[0128] Here, we have described a method for generating content elements and adding context information from content based on conventional formats that are already being distributed and delivered, but based on the mechanism proposed in this technology, it is also possible to directly create works that correspond to content elements.
[0129] (Example of the scenario generation tool UI) Here, the user interface of a scenario generation tool for generating a scenario will be described with reference to Figures 14 to 17. This scenario generation tool is executed by the control unit 200 of the editing device 20 operated by a producer or the like, and various screens are displayed on the display 231.
[0130] When the scenario generation tool is started, the scenario selection / new creation screen is displayed as shown in Fig. 14. This scenario selection / new creation screen includes a map / scenario display area 251, a scenario list 252, and a new scenario creation button 253.
[0131] The names of the scenarios are displayed on pins 261A that indicate positions on the map in the map / scenario display area 251, or the scenario display banners 262A are displayed as a list in a predetermined order such as alphabetical order in the scenario list 252. The create new scenario button 253 is operated when creating a new scenario.
[0132] The creator can select a desired scenario by clicking on a pin 261A on the map corresponding to the desired area or on a scenario display banner 262A in the scenario list 252.
[0133] At this time, if attention is focused on pin 261B among the multiple pins 261A, the name of the scenario corresponding to pin 261B, "Scenario #1," is displayed in a balloon because it has been selected by cursor 260. Then, if edit button 262B is clicked while scenario #1 corresponding to pin 261B is selected, the scenario edit screen of Fig. 15 is displayed.
[0134] The scenario editing screen of FIG. 15 includes a map / geofence display area 254, a geofence list 255, and an editing tool display area 256.
[0135] The names of the geofences are displayed in the geofence areas 271A to 271E that represent the geofence areas on the map in the map / geofence display area 254, or the geofence display banner 272A is displayed as a list in a predetermined order, such as alphabetical order, in the geofence list 255.
[0136] The geofence areas 271A to 271E can be set to various shapes such as a circle or a polygon.
[0137] In the map / geofence display area 254, context information assigned to the activation conditions (activation range) for which default values are set is displayed as text within each geofence, or in a speech bubble when the desired geofence is selected. Based on this display, the creator can check the context information associated with the activation range of each content element.
[0138] This allows the creator to select a desired geofence by clicking on the geofence areas 271A to 271E on the map corresponding to the desired area or on the geofence display banner 272A in the geofence list 255.
[0139] The editing tool display area 256 includes a circular geofence creation button 273, a polygonal geofence creation button 274, a geofence movement button 275, an overwrite save button 276, a new save button 277, a delete button 278, and a back button 279.
[0140] The circular geofence creation button 273 is operated when creating a geofence having a circular shape. The polygonal geofence creation button 274 is operated when creating a geofence having a polygonal shape. The geofence movement button 275 is operated when moving a desired geofence.
[0141] The overwrite button 276 is operated when saving the scenario to be edited by overwriting an existing scenario. The new save button 277 is operated when saving the scenario to be edited as a new scenario. The delete button 278 is operated when deleting the scenario to be edited. The back button 279 is operated when returning to the scenario selection / new creation screen.
[0142] Here, if we focus on the patterned geofence area 271C among the geofence areas 271A to 271E, it is selected by the cursor 260, so the geofence name corresponding to the geofence area 271C, which is "Geofence #1", is displayed in a speech bubble, and the content element set in the geofence may be played.
[0143] Then, when the edit button 272B is clicked while the geofence #1 corresponding to the geofence area 271C is selected, the geofence edit screen of FIG. 16 is displayed.
[0144] 16 includes a geofence detailed setting area 257. The geofence detailed setting area 257 includes detailed setting items for the geofence, such as the geofence name, center position, radius, playback time, weather, content elements, playback range, volume, repeat playback, fade-in / out, and playback priority level.
[0145] The geofence name corresponds to the context setting item. The center position, radius, playback time, and weather correspond to the activation condition setting items, and their default values are set here. The content element, playback range, volume, repeat playback, fade-in / out, and playback priority level correspond to the content element and playback condition setting items, and their default values are set here.
[0146] In the geofence name input field 281A, "geofence #1" is input as the geofence name.
[0147] In the center position input field 281B and the radius input field 281C, "latitude, longitude" and "80 m" are input as default values for the center position and radius of the circular geofence.
[0148] In the playback time input field 281D, "7:00 - 10:00" is input as the default value for the playback time. In addition, since the weather input field 281E is set to "Not specified", the default value for the weather is not set.
[0149] In the content element input field 281F, "http:xxx.com / sound / folder#1 / 01.mp3" is entered as the default value of the content element. This can be entered using a content element selection screen 283 that is displayed by clicking on the selection button 282.
[0150] The content element selection screen 283 displays the audio file data of the content elements stored in the storage unit 103 of the data management server 10. In this example, by selecting a desired folder from the folders displayed in a hierarchical structure on the content element selection screen 283, a desired audio file within the folder can be selected.
[0151] Here, a search process may be performed using the desired keyword input in the search keyword input field 284A as a search condition, and a list of desired audio files according to the search results may be presented.
[0152] In the playback range input field 281G and the volume input field 281H, "00:00:08 - 00:01:35" and "5" are input as default values for the playback range and volume. Note that the playback time and volume may be input automatically depending on the content element.
[0153] In the repeat playback input field 281I and the fade-in / out input field 281J, "Repeat playback: yes" and "Fade-in / out: yes" are input as default values for repeat playback and fade-in / fade-out of the audio file.
[0154] In the playback priority level input field 281K, "1" is input as the default value of the playback priority level. The playback priority level can be set in predetermined stages such as three stages from "1" to "3" or five stages from "1" to "5", with the lower the number the higher the priority and the higher the number the lower the priority.
[0155] Note that the geofence editing screen in FIG. 16 shows the case where the shape of geofence #1 is circular, but if the shape is polygonal (rectangle), the geofence editing screen in FIG. 17 is displayed.
[0156] The geofence editing screen in Figure 17 differs from the geofence editing screen shown in Figure 16 in that the setting items for the activation conditions are the vertex positions of a rectangular geofence instead of the center position and radius of a circular geofence.
[0157] 17, a vertex position input field 291B consisting of a list box is provided instead of the text boxes of the center position input field 281B and the radius input field 281C in FIG.
[0158] In this example, the vertex position input field 291B displays a list of multiple latitude and longitude combinations, such as latitude #1 and longitude #1, latitude #2 and longitude #2, latitude #3 and longitude #3, etc., and the desired latitude and longitude combination selected from the list is set as the default value for the vertex position of the rectangular geofence.
[0159] The above-described user interface of the scenario generation tool is an example, and other user interfaces may be used, such as using other widgets instead of text boxes and radio buttons.
[0160] For example, on the geofence editing screen, a drop-down list or a combo box can be used instead of the text boxes that make up the playback time input field 281D, the weather input field 281E, the volume input field 281H, or the playback priority level input field 281K, or the list box that makes up the vertex position input field 291B.
[0161] (Overall processing) Next, an overview of information processing in the second embodiment will be described with reference to FIG.
[0162] 18 is realized by at least cooperation between the data management server 10 (control unit 100 thereof) and the editing device 20 (control unit 200 thereof) in the information processing system 1. That is, this information processing is executed by at least one of the control units 100 and 200.
[0163] As shown in FIG. 18, in the information processing system 1, one or more content elements (e.g., "character dialogue") consisting of at least some of the media are extracted from content (movies, animations, games, etc.) consisting of multiple media (video, audio, etc.) (S201), and a context (e.g., a context in which the dialogue is expected to be heard) is generated for the content element (S202).
[0164] Then, in the information processing system 1, context information (for example, "gaining courage") is assigned to each content element (for example, "character's lines") (S203). As a result, the content element and the context information are stored in association with each other in the content element-context information DB 151.
[0165] Furthermore, one or more data sets of "content element-context information" are stored as a scenario (for example, "starting town") in the scenario DB 152 (S204). Here, the data sets can be packaged based on a certain theme (such as the title of the work that was the source of the re-edit, the setting, the emotions evoked, etc.) and stored in the scenario DB 152 (S211).
[0166] Here, the content element may include, for example, a part (such as a part of a song) of streaming content (such as a song distributed by a music streaming service). In this case, in order to identify the part of the streaming content, the content ID and playback range of the content may be specified (S221), and information indicating the content ID and playback range may be stored in the content element-context information DB 151 in association with the target context information.
[0167] Furthermore, introductory content (other content elements) such as characters may be generated for the content elements (S231), and the introductory content may be presented before the content elements are played. For example, before playing a song (content element) distributed from a music streaming service, an introductory statement may be presented by a specific voice character (e.g., a disc jockey (DJ) character) corresponding to the context information.
[0168] Furthermore, by performing machine learning on the relationship between the content elements and the context information stored in the content element-context information DB 151 (S241), it is possible to automatically assign context information to new content elements.
[0169] Here, various techniques such as neural networks (NN) can be used as machine learning techniques. For example, a technique for identifying elements such as people, living things, objects, buildings, and scenery contained in a scene from image information or audio information contained in a video or still image can be used to determine the range of content elements and automatically generate one or more pieces of context information assumed from the identification results or a combination thereof.
[0170] The second embodiment has been described above.
[0171] <3. Third Embodiment>
[0172] Incidentally, when generating a combination of content elements and context information from content consisting only of text, such as an e-book novel, the extracted text itself can be used as a content element, and can be displayed, for example, as a character image on a display device such as a public display or AR glasses. However, audio (sound) can also be used. AR glasses are eyeglass-type devices that support AR (Augmented Reality).
[0173] That is, voice data can be generated from text data used as a content element using TTS (Text To Speech) technology, and the voice data can be used as a content element.
[0174] In addition, machine learning techniques may be used to search for or synthesize data such as audio data or image data that has a related impression (image) from text that makes up words or sentences, and use the data as content elements.
[0175] On the other hand, for content consisting only of audio data or image data, machine learning techniques can be used to search for or synthesize text that constitutes related words or sentences, and then use that text as a content element. In other words, this allows for adding content not included in the existing content, or adding expressions in other modalities not included in the original content, such as tactile sensations.
[0176] Note that TTS technology is an example of a speech synthesis technology that artificially creates a human voice, and other technologies may be used to generate the voice. Alternatively, a recording of a human reading may be used. Also, while the above explanation shows a case where machine learning technology is used, data as content elements may also be generated by separately analyzing acquired data.
[0177] (Overall processing) Next, an overview of information processing in the third embodiment will be described with reference to FIG.
[0178] The information processing shown in FIG. 19 is realized by at least cooperation between the data management server 10 (control unit 100 thereof) and the editing device 20 (control unit 200 thereof) in the information processing system 1.
[0179] As shown in FIG. 19, in the information processing system 1, one or more content elements (e.g., a sentence in a novel) consisting of a first medium (e.g., text) are extracted from content (e.g., an e-book novel) consisting of multiple media (e.g., text) (S301), and a content element (e.g., a voice corresponding to a sentence in the novel) consisting of a second medium (e.g., TTS voice) is generated (S302).
[0180] Then, in the information processing system 1, context information (for example, information about the context in which the audio of the sentence in the novel is expected to be heard) is assigned to each content element (for example, audio corresponding to a sentence in a novel) (S303), and the content element and the context information are associated and stored in the content element-context information DB151.
[0181] Furthermore, one or more data sets of "content element-context information" are saved (accumulated) as a scenario in the scenario DB 152 (S304).
[0182] Here, by performing machine learning in advance to determine the relationship between a first medium (such as text) and a second medium (such as TTS voice) (S311), content elements of the second medium can be generated from content elements of the first medium based on the results of that machine learning.
[0183] The third embodiment has been described above.
[0184] <4. Fourth embodiment>
[0185] By using the user scenario generation tool, the user can obtain a desired scenario and a desired data set of "content element-context information" on the playback device 30 that the user owns.
[0186] In other words, by executing the user scenario generation tool on the playback device 30, multiple "content element-context information" data sets contained in the acquired scenario can be displayed, and activation conditions consisting of a combination of senseable conditions can be set for each "content element-context information" data set using a user interface for placing them in the actual space around the user.
[0187] The conditions for this activation can include, for example, information about the Global Positioning System (GPS), location information such as latitude and longitude estimated from information from wireless LAN (Local Area Network) access points, and usage and authentication information obtained from wireless beacons and short-range wireless communication history.
[0188] Furthermore, activation conditions include, for example, information about the user's position, posture, behavior, and surrounding environment estimated from images captured by a camera, information about the time and duration measured by an environmental information clock, environmental information and authentication information based on audio information obtained from a microphone, information about the body's posture, movement, riding status, etc. obtained from an inertial sensor, and information about the breathing rate, pulse rate, emotions, etc. estimated from biosignal information.
[0189] For example, as shown in Figure 20, if the text "gaining courage" is attached to audio content extracted from the lines of a certain character as a data set of "content element-context information," the "latitude and longitude" estimated from GPS information, etc., can be set as the activation condition.
[0190] The activation conditions can be set using a user scenario generation tool, but they can also be completed before using the service, or the tool can be launched and set while using the service.
[0191] Here, as an example of a user scenario generation tool, we will explain a case where a data set of "content elements - context information" is displayed on a map, and the user uses an interface placed on the map to set the range and time period on the map as sensing activation conditions.
[0192] A user can create a desired user scenario by operating a user scenario generation tool executed on a playback device 30 such as a smartphone or an information device such as a personal computer. The user scenario generation tool may be provided as a native application or as a web application using a browser.
[0193] (Example of the user scenario generation tool UI) 21 to 25, the user interface of a user scenario generation tool executed by a playback device 30 such as a smartphone will be described. This user scenario generation tool is executed by, for example, the control unit 300 of the playback device 30 operated by a user, and various screens are displayed on the display 331.
[0194] When the user scenario generation tool is started, the scenario selection and playback screen is displayed as shown in Fig. 21. This scenario selection and playback screen includes a map and scenario display area 411, a scenario list 412, and a new scenario creation button 413.
[0195] The scenarios are displayed in a map / scenario display area 411 by name on a pin 411A that indicates a location on the map, or in a scenario list 412 as a list in a predetermined order such as alphabetical order or order of shortest distance from the current location.
[0196] To create a new user scenario, simply tap the Create New Scenario button 413. The scenario selection / playback screen may also perform a search using the desired keyword entered in the search keyword input field 414 as a search condition, and present a scenario based on the search results.
[0197] The user can select a desired scenario by tapping a pin 411A on the map corresponding to a desired area or a scenario display banner 412A in the scenario list 412.
[0198] In this example, scenario #1 is currently playing, and scenarios #2 and #3 are stopped, among the scenario display banners 412A displayed in the scenario list 412. Note that in this example, only three scenario display banners 412A are displayed, but other scenarios may also be displayed by flicking the screen to scroll, for example.
[0199] At this time, if one focuses on pin 411B among multiple pins 411A in map / scenario display area 411, pin 411B is in a selected state, and therefore the scenario name corresponding to pin 411B, which is "Scenario #1," is displayed in a balloon. If edit button 412B is tapped while scenario #1 corresponding to pin 411B is selected, the activation condition setting screen of Fig. 22 is displayed as the scenario editing screen.
[0200] The activation condition setting screen of FIG. 22 includes a map / geofence display area 421 , an overwrite save button 422 , a new save button 423 , a delete button 424 , and a back button 425 .
[0201] Geofence areas 421A to 421E are displayed on a map of a desired area in the map / geofence display area 421. The geofence areas 421A to 421E can be set to various shapes such as a circle or a polygon.
[0202] In the map / geofence display area 421, context information assigned to the activation conditions (activation range) is displayed as text within each geofence, or in a speech bubble when the desired geofence is tapped. Based on this display, the user can check the context information associated with the activation range of each content element.
[0203] The geofence can be moved on the screen. Here, if you focus on the patterned geofence area 421C among the geofence areas 421A to 421E, the geofence name corresponding to the geofence area 421C, which is "geofence #1", is displayed in a balloon because it is in a selected state.
[0204] Here, the user uses finger 400 to select geofence area 421C, and moves it diagonally downward to the right (in the direction of the arrow in the drawing) to move the position.
[0205] Also, although not shown in the figure, when geofence area 421C is selected, the area of geofence area 421C may be enlarged or reduced by performing a pinch-out operation or a pinch-in operation, or the shape of geofence area 421C may be deformed in accordance with a specified operation.
[0206] To save the settings of this activation condition as scenario #1, tap the overwrite save button 422, while to save it as a new scenario, tap the new save button 423. The delete button 424 is operated to delete scenario #1. The back button 425 is operated to return to the scenario selection / playback screen.
[0207] Furthermore, when the user performs a long press operation on geofence area 421C using finger 400, the activation condition detailed setting screen of FIG. 23 is displayed.
[0208] The activation condition detailed setting screen of FIG. 23 includes a geofence detailed setting area 431, a save button 432, and a back button 433.
[0209] The geofence detailed setting area 431 includes a geofence name input field 431A, a center position input field 431B, a radius input field 431C, a playback time input field 431D, a weather input field 431E, a content element input field 431F, a playback range input field 431G, a volume input field 431H, a repeat playback input field 431I, a fade-in / out input field 431J, and a playback priority level input field 431K.
[0210] The geofence name input field 431A to the playback priority level input field 431K correspond to the geofence name input field 281A to the playback priority level input field 281K in FIG. 16, and the values set there as default values are displayed as they are.
[0211] The save button 432 is operated to save the settings of the geofence # 1. The return button 433 is operated to return to the activation condition setting screen.
[0212] The user may use the default settings of geofence #1 as they are, or may change them to desired settings. For example, when the content element input field 431F is tapped, the content element selection screen of FIG. 24 is displayed.
[0213] The content element selection screen of FIG. 24 includes a content element display area 441, a selection button 442, and a back button 443.
[0214] In the content element display area 441, icons 441A to 441F corresponding to the respective content elements are arranged in a tiled pattern of three rows and two columns.
[0215] The select button 442 is operated to select a desired icon from the icons 441 A to 441 F. The return button 443 is operated to return to the activation condition detailed setting screen.
[0216] Here, when the user uses finger 400 to perform a tap operation on icon 441A among icons 441A to 441F, content element #1 is played back.
[0217] Furthermore, when the user performs a long press on selected icon 441A using finger 400, the content element editing screen of FIG. 25 is displayed.
[0218] The content element edit screen of FIG. 25 includes a content playback portion display area 451, a content playback operation area 452, a song change button 453, and a back button 454.
[0219] In order to edit content element #1 as a song, content playback portion display area 451 displays the waveform of the song of content element #1, and you can specify the portion you want to play by sliding sliders 451a and 451b left and right.
[0220] In this example, among the waveforms of the music of content element #1, the waveform of the music in cut selection area 451B corresponding to the area outside sliders 451a and 451b is set as the waveform not to be played, and the waveform of the music in playback selection area 451A corresponding to the area inside sliders 451a and 451b is set as the waveform to be played. Note that seek bar 451c indicates the playback position of the music of content element #1 being played.
[0221] In the content playback operation area 452, buttons for operating the music piece of the content element #1, such as a play button, a stop button, and a skip button, are displayed.
[0222] The user can check the waveform of the song in the content playback portion display area 451 and operate the buttons and sliders 451a, 451b, etc. in the content playback operation area 452 to extract only the portion of the song in content element #1 that they want to play.
[0223] The change song button 453 is operated when changing the song to be edited, and the return button 454 is operated when returning to the activation condition detailed setting screen.
[0224] In this way, the user can create a desired user scenario by operating the user scenario generation tool executed by the playback device 30 such as a smartphone.
[0225] Next, the user interface of the user scenario generation tool executed by an information device such as a personal computer will be described with reference to FIGS.
[0226] When the user scenario generation tool is started, the scenario selection screen shown in Fig. 26 is displayed. This scenario selection screen includes a map / scenario display area 471 and a scenario list 472.
[0227] The names of the scenarios are displayed on pins 471A that indicate positions on the map in a map / scenario display area 471, or scenario display banners 472A are displayed in a scenario list 472 in a predetermined order as a list.
[0228] The user can select a desired scenario by clicking a pin 471A on the desired map or a scenario display banner 472A in the scenario list 472.
[0229] When the edit button 472B is clicked, a scenario edit screen for editing the scenario is displayed. When a new scenario is to be created, a new scenario creation button (not shown) is operated.
[0230] When the user selects a desired scenario, the activation condition setting screen shown in Fig. 27 is displayed. This activation condition setting screen includes a map / geofence display area 481 and a context list 482.
[0231] A geofence area 481A indicating the activation range of a content element is displayed in the map / geofence display area 481. The geofence area 481A is represented by a plurality of preset shapes such as circles or polygons.
[0232] In the map / geofence display area 481, the context information assigned to the activation conditions (activation range) is displayed as text within the geofence area 481A, or displayed in a speech bubble when the desired geofence area 481A is clicked.
[0233] The geofence area 481A can be moved on the screen in response to a drag operation. Here, if attention is focused on the patterned geofence area 481B among the multiple geofence areas 481A, the geofence area 481B can be moved diagonally upward to the right (in the direction of the arrow in FIG. 28) by a drag operation, and can be moved from the position shown in FIG. 27 to the position shown in FIG. 28.
[0234] In addition, by placing the cursor on the white circle (◯) on the thick line that indicates the shape of the geofence area 481B and dragging it in a desired direction, the shape of the geofence area 481B can be transformed into a desired shape.
[0235] In this way, the user can set for himself or herself where in the real-life space the context corresponds by moving or transforming the geofence area 481B based on the context information displayed in the geofence area 481B.
[0236] The content elements may be presented in the form of a separate list. Furthermore, content elements that are not used may be deleted, and separately obtained content elements may be added to the scenario currently being edited.
[0237] Here, when the edit button 482B of the context display banner 482A corresponding to the geofence area 481B in the context list 482 is clicked or a predetermined operation is performed on the geofence area 481B, the geofence edit screen of Figure 29 is displayed.
[0238] This geofence editing screen includes a geofence detail setting area 491, a select button 492, an update button 493, a delete button 494, and a cancel button 495.
[0239] The geofence detail setting area 491 includes a geofence name input field 491A, a content element input field 491B, a repeat playback input field 491C, a fade-in / out input field 491D, a playback range input field 491E, and a volume input field 491F. These setting items correspond to the setting items in the geofence detail setting area 431 in FIG. 23.
[0240] Furthermore, when the select button 492 is clicked, a desired content element can be selected using the content element selection screen, similar to the select button 282 in Fig. 16. The update button 493 is operated when updating the setting items of the geofence area 481B. The delete button 494 is operated when deleting the geofence area 481B. The cancel button 495 is operated when canceling editing.
[0241] In this way, a user can create a desired user scenario by operating a user scenario generation tool executed by an information device such as a personal computer.
[0242] In the above explanation, a user interface using a map was used as an example of a user scenario generation tool, but other user interfaces that do not use a map may also be used. Below, we will explain a method for setting activation conditions without using a map.
[0243] For example, if you want to set an object that is not shown on the map, such as a "bench in the square in front of the station," to be activated around that object, you can set it by taking a picture of the desired bench with the camera unit 306 of the playback device 30, such as a smartphone.
[0244] Alternatively, the user can take a picture of the desired bench with the camera of the wearable device they are wearing and issue a voice command such as "Take a picture here" or "Set it on this bench" to set it. Furthermore, if the user can take a picture of their own hands using a camera such as eyewear, they can make a hand gesture to surround the bench and record the objects and scenery within the surround when the gesture is recognized.
[0245] Furthermore, even when setting activation conditions that cannot be set using map representations, such as the user's biological state or emotions, a "current feeling" button can be displayed on the playback device 30 such as a smartphone, and data and recognition results can be recorded at the time the button is tapped or clicked, or for a certain period of time before and after, and set as the activation condition. Note that, as in the above case, input can also be made using the user's voice, gesture commands, etc.
[0246] Here, in order to easily set multiple pieces of data, for example, a "Current Situation" button may be displayed, or may be preset as a voice command or a specific gesture, and when input is made to the button, data such as the pre-specified location, time, weather, surrounding objects, weather, biometric data, and emotions may be acquired all at once.
[0247] By providing these input methods, especially input methods that do not involve a screen, users can easily input data in their daily lives while experiencing a service or while the service is stopped.
[0248] In this way, data entered by the user without using a screen is transmitted to, for example, the data management server 10 and stored in the user scenario DB 153. This allows the user to display the screen of the user scenario generation tool on the playback device 30 that the user owns. The user can then check and re-edit the link between the activation conditions displayed on this screen and the "content element-context information" data set.
[0249] The above operations are operations in which the user sets only the activation conditions for content elements in the scenario provided, but depending on the terms of use, the content of the content elements, such as audio data and image data, or the context information attached to the content elements may also be permitted as operations that the user can change.
[0250] The scenario after editing is stored as a user scenario in the user scenario DB 153. Note that the user scenario stored in the user scenario DB 153 can also be made available to other users using a sharing means such as a social networking service (SNS).
[0251] Furthermore, by displaying the data sets of multiple "content element-context information" contained in a scenario in an editing tool such as a user scenario generation tool, and allowing the user to link them to the actual location, time of day, environment, and their own actions and emotions in their living space, this can be applied to services such as the following:
[0252] In other words, as an example of one service, consider the case where a scenario is acquired that consists of multiple data sets of "content elements-context information" that are made up of lines spoken by a specific character appearing in an anime work in various contexts.
[0253] In this case, while referring to presented context information such as "home," "station," "street," "intersection," "cafe," and "convenience store," the user subjectively inputs the locations of "home," "station," "street," "intersection," "cafe," and "convenience store" where the user actually lives as activation conditions using editing means such as a user scenario generation tool. This allows the user to receive playback of content elements according to the context on the playback device 30 they own in the place where they live and in a place with the context they envision (for example, an intersection).
[0254] FIG. 30 shows an example of a user scenario setting.
[0255] In FIG. 30, two users, User A and User B, set activation conditions A and B, respectively, for a scenario to be distributed, and each creates his or her own user scenario.
[0256] In this case, when setting the activation conditions for the same scenario, user A sets activation condition A and user B sets activation condition B, so the activation conditions differ for each user.
[0257] Therefore, the same scenario can be executed by different users in different locations. In other words, one scenario can be used by users living in different locations.
[0258] Another example of a service involves collaboration with streaming distribution services.
[0259] For example, conventional music streaming services create and distribute playlists that compile audio data from multiple works in existing music formats (such as single songs) based on a certain theme, such as a specific creator or usage scenario.
[0260] In contrast, with this technology, the work itself, or a part of the work that expresses a specific context, is extracted and made into a content element, and context information that describes the situation (for example, a train station at dusk) or state (for example, a tired person on the way home) in which the music is played is added to the content element, and this is compiled into a scenario, stored in the scenario DB 152, and made available for distribution.
[0261] The user can obtain the above scenario using the playback device 30, and create a user scenario by placing the contained multiple ``content element-context information'' data sets at specific locations and time periods in their own living area while referring to the attached context information, and register the scenario in the user scenario DB 153.
[0262] When editing a user scenario, the user can specify a portion of the work that they want to play as a content element by specifying the playback range. A scenario can include content elements (other content elements) as voice characters that explain the work being played when or between content element playback.
[0263] It should be noted that this voice character can be obtained not only through the same route as the scenario, but also through a route different from the scenario, and for example, the user can have the character they prefer from among multiple voice characters provide the explanation.
[0264] The scenario DB 152 stores combinations of context information for various content elements that are intended to be provided to users by creators.
[0265] For example, if a recognizer that uses this context information as training data and machine-learns the melodic structure of content elements is used, it can estimate the context that is likely to be evoked from the melodic structure of a content element in a way that reflects the creator's subjective tendencies. This estimation result can then be used to automate the process of assigning context information to content elements, or to support creators in assigning context information by presenting multiple contexts with a certain degree of correlation.
[0266] In addition, the user scenario DB 153 sequentially stores data sets of "content element-context information" linked by the user to activation conditions consisting of the location, time, environment, physical state, emotion, etc. of the user's living space.
[0267] In other words, the user scenario DB153 stores a large number of data sets of "content element-context information" for which activation conditions have been set by multiple users, and by using machine learning or analyzing this stored information, it is possible to create algorithms or recognizers that automate processes.
[0268] Furthermore, for example, it is possible to analyze trends in the context information assigned to a real-world (real space) location with a specific latitude and longitude from information about multiple users stored in the user scenario DB 153.
[0269] For example, if it is analyzed that a park at the exit of a certain real-world station tends to be set as an "uplifting" or similar context, the results of that analysis can be used to utilize the data for other services, such as selling food or books that are expected to uplift people in that park.
[0270] Also, for example, if a specific context is set for the scenery seen from a certain location at a certain time of day, linking a content element of a work, such as a phrase in a song, to the lyrics, this information can be fed back to the composer or lyricist of the song and used as reference data when creating subsequent works.
[0271] (Overall processing) Next, an overview of information processing in the fourth embodiment will be described with reference to FIGS.
[0272] 31 and 32 is realized by at least cooperation between the data management server 10 (control unit 100 thereof) and the playback device 30 (control unit 300 thereof) in the information processing system 1. That is, this information processing is executed by at least one of the control units 100 and 300.
[0273] As shown in FIG. 31, in the information processing system 1, context information is assigned to each content element, and one or more data sets of "content element-context information" are stored as a scenario in the scenario DB 152 (S401).
[0274] At this time, in the information processing system 1, for each piece of context information assigned to the content element, an activation condition according to sensor data obtained by sensing the user is set (S402). As a result, a user scenario consisting of a data set of context information and activation conditions specific to the user is generated (S403), and is stored in the user scenario DB 153 (S404).
[0275] Here, the activation conditions can be set according to the captured image data, characteristic operation data, etc. Here, the image data includes data on the image that is assumed to be viewed by the user. Furthermore, the characteristic operation data includes, for example, data on the operation of a button (current feeling button) for registering information according to the user's current emotion.
[0276] In addition, by performing machine learning on the relationship between the context information (such as "gaining courage") stored in the user scenario DB153 and the activation conditions (such as the exit of a specific station) (S411), the results of the machine learning can be output.
[0277] More specifically, context information can be automatically generated for a specific activation condition according to the results of machine learning (S421). For example, if the results of machine learning identify that a location according to sensor data is a place where courage can be gained, "gain courage" is generated as context information and assigned to the target content element.
[0278] Furthermore, depending on the results of machine learning, it is possible to automatically generate an activation condition corresponding to the user for specific context information (S431). For example, if the learning result identifies that the place where courage is gained is in the vicinity of the user, location information corresponding to the location is set as an activation condition for the context information "gain courage."
[0279] 32, the information processing system 1 provides a user scenario generation tool as a user interface using a map for setting user-specific activation conditions. As mentioned above, this user scenario generation tool is provided as an application executed by a playback device 30 such as a smartphone or an information device such as a personal computer.
[0280] In the information processing system 1, an activation condition is set for each piece of context information assigned to a content element extracted from content (S401, S402).
[0281] Here, by utilizing a user scenario generation tool, an interface is provided that allows a data set of content elements and context information to be presented on a map of a desired area (S441), and a specified area to be set on the map of the desired area as an activation condition for the context information (S442).
[0282] The fourth embodiment has been described above.
[0283] <5. Fifth Embodiment>
[0284] In the information processing system 1, sensor data such as the user's location, physical state, emotions, movements, information on objects, structures, buildings, products, people, animals, etc. in the surrounding environment, and the current time are sequentially acquired by sensing means implemented in a playback device 30 held or worn by a user, or in a device located in the user's vicinity.
[0285] Then, the determining means sequentially determines whether or not these data or a combination of data matches the activation conditions set by the user.
[0286] Here, if it is determined that the activation condition matches the sensor data from the sensing means, the content elements included in the "content element-context information" data set linked to the activation condition are played from a pre-specified device (e.g., playback device 30) or a combination of multiple devices (e.g., playback device 30 and devices located in the vicinity).
[0287] Here, the location and timing of playback are determined by comparing sensor data from the sensing means with the activation conditions, and the judgment process does not directly include subjective elements such as context or a machine learning recognizer that uses data that includes subjective elements, enabling the system to operate reproducibly and stably.
[0288] On the other hand, since the user is the one who actively combines the activation conditions with the ``content element-context information'' data set, there is also the advantage that it is easier for the user to understand that the content element is being presented in an appropriate situation.
[0289] FIG. 33 shows examples of combinations of activation conditions and sensing means.
[0290] As a temporal activation condition, it is possible to set the time of day or duration, and it is possible to measure and determine using a clock, timer, etc. Furthermore, as a spatial activation condition, it is possible to set a position, such as latitude, longitude, or approach to a specific location, and it is possible to measure and determine using GPS, Wi-Fi (registered trademark), wireless beacon, etc.
[0291] Also, authentication information such as a user ID may be set as an activation condition, and it is possible to measure and determine using proximity communication such as Bluetooth (registered trademark), etc. Furthermore, the user's posture (standing, sitting, lying down, etc.) or the user's behavior (on a train, bicycle, escalator, etc.) may also be set as an activation condition, and it is possible to measure and determine using an inertial sensor, a camera, proximity communication, etc.
[0292] Additionally, surrounding environmental information such as chairs, desks, trees, buildings, rooms, scenery, and scenes may be set as activation conditions, and can be measured and determined using cameras, RF tags, wireless beacons, ultrasound, etc. Furthermore, physical posture, movement, respiratory rate, pulse rate, emotions, and other states may be set as activation conditions, and can be measured and determined using inertial sensors, biosensors, etc.
[0293] The combinations shown in the table of FIG. 33 are merely examples, and the activation conditions and sensing means are not limited to those shown in this table.
[0294] The fifth embodiment has been described above.
[0295] 6. Sixth Embodiment
[0296] However, it is possible that the same activation condition is set for two or more content elements included in at least one scenario. For example, in a data set of multiple content elements and content information in which activation conditions are set within a certain range on a map, two or more activation ranges may be set overlapping and include the same location on the map.
[0297] Specifically, as shown in FIG. 34, on a map 651, a geofence 661 set as a circular activation range and geofences 662A to 662E set as circular activation ranges within that circle are superimposed.
[0298] In this case, when content elements are played back on the playback device 30, for example, if all content elements are played back simultaneously according to a pre-set rule, it is possible that not all content elements will be played back when some content elements are played back based on the set priority.
[0299] Here, by preparing in advance a user scenario for setting the presentation range, which is referenced when the triggering condition is satisfied in the user scenario, it is possible to appropriately play back the content element.
[0300] Specifically, as shown in Figure 35, an example is shown in which the content element is the reading of a sentence by TTS voice, and the user scenario for setting the presentation range specifies an utterance (line) by character A for activation condition A, which includes the entire activation range including the home, etc., and an utterance (line) by character B for activation condition B, which also includes the activation range including the home, etc.
[0301] 35, the lower layer L1 corresponds to a user scenario, and the upper layer L2 corresponds to a user scenario for setting a presentation range. In the lower layer L1, an elliptical area corresponds to an activation range set by a geofence.
[0302] In this case, if the activation conditions for the character's activity range setting scenario are set to exclusive, character B will speak when activation condition C1 of the user scenario is met, and character A will speak when activation condition C2 is met. In other words, in this case, there will always be one character.
[0303] On the other hand, if the activation conditions for the character activity range setting scenario are not exclusive, when the user scenario activation condition C1 is met, either character A or B will speak. Which of character A or B will speak may be determined randomly, or a specific rule may be set. Also, when activation condition C2 is met, only character A will speak. In other words, in this case, when the user is at home, there will be two characters.
[0304] In addition, the priority order can be set based on sensor data. For example, when multiple content elements are speeches (lines) by multiple characters, it is assumed that all of the corresponding content elements are in a playable state when the user is positioned at a location where the activation conditions of the multiple content elements overlap.
[0305] At this time, as shown in FIG. 36, based on the relative positional relationship between the position of the user 600 and specific positions 671A to 671C (e.g., the center of the circle) in the activation range of the content elements corresponding to the geofences 672A to 672C, and the sensor data in the direction in front of the body of the user 600 (e.g., the upper right direction in the figure), only the content elements of the geofence 672A located in front of the body are played.
[0306] At this time, if the user 600 is wearing stereo earphones connected to the playback device 30, the fixed position of the sound source (e.g., dialogue) being played can be controlled three-dimensionally (sound image localization) according to the relative positional relationship between the position of the user 600 and specific positions 671A to 671C in the activation range of the content element according to the geofences 672A to 672C.
[0307] By using the above control, it is possible to obtain playback of the character's speech in the direction that the user 600 is facing, making it possible to select the presentation of a sound source (e.g., lines) by a desired character depending on the orientation of the user 600's body, head, etc.
[0308] 37, the volume of the sound source produced by the character may be changed depending on the position of the user 600 in the geofence 672A. For example, the volume of the sound source may be increased as the user 600 approaches a specific position 671A, and decreased as the user 600 moves away from the specific position 671A.
[0309] Furthermore, by relating the acceptance of a spoken command from the user 600 to the activation condition, it is possible to realize a guidance service in which, when the user 600 faces a certain direction and asks a question, a character set in that direction will present information related to that position.
[0310] Here too, the user scenario for setting the presentation range may be referenced.
[0311] Specifically, as shown in Fig. 38, the user scenario for setting the presentation range includes information specifying the sound source setting positions P1 to P4 as well as information setting the activation range for each of the activation conditions C1 to C4. However, the sound source setting positions P1 to P4 are not limited to positions within the activation range that specifies the activation conditions C1 to C4.
[0312] Figure 38 shows four activation conditions C1 to C4 that have a common activation condition area CA (diagonal lines in the figure), and each activation condition C1 to C4 has a sound source setting position P1 to P4 (black circle in the figure) set.
[0313] At this time, if the activation conditions are satisfied in the user scenario, that is, if the user 600 enters the common activation condition area CA, the sound source setting positions are searched for for all the activation conditions whose conditions are satisfied.
[0314] Here, among the searched sound source setting positions P1 to P4, the sound source setting position P2 within the viewing angle region VA calculated from the user orientation information measured by the sensor unit 305 of the playback device 30 carried by the user 600 is identified. Then, the content element associated with the activation condition C2 having the identified sound source setting position P2 is played back.
[0315] The above-described control is an example of control when two or more activation ranges are set to overlap and include the same location on the map, and other control may be performed. For example, when all content elements are played simultaneously, control may be performed such that one content element is background music and the other content elements are multiple lines of dialogue, thereby presenting an expression in which multiple lines of dialogue are played against the same background music as the user moves within the activation range.
[0316] (Multiple character arrangement) Furthermore, the above-described control is not limited to the presentation of audio (sound), but can also be used to similarly control the presentation of character images through a display device such as an eyeglass-type device compatible with augmented reality (AR). Next, a case where the placement of multiple characters can be set for a scenario will be described with reference to Figures 39 to 45.
[0317] FIG. 39 shows an example of the configuration of the information processing system 1 in the case where the arrangement of multiple characters can be set.
[0318] 39 shows the data management server 10 and the playback device 30, which are among the devices that make up the information processing system 1 in FIG. 2. However, some of the processes executed by the data management server 10 may be executed by other devices, such as the editing device 20 or the playback device 30.
[0319] In the playback device 30, the control unit 300 includes a user position detection unit 341, a user direction detection unit 342, a voice recognition intent understanding unit 343, and a content playback unit 344.
[0320] The user position detection unit 341 detects the user's position based on information related to GPS and the like.
[0321] The user direction detection unit 342 detects the direction in which the user is facing based on sensor data from the sensor unit 305 (FIG. 5).
[0322] The voice recognition and intention understanding unit 343 performs voice recognition and intention understanding processing based on the voice data of the user's utterance, and understands the intention of the user's utterance.
[0323] It should be noted that this voice recognition and intent understanding process is not limited to being performed by the control unit 300, and a part or all of the process may be performed by a server on the Internet 40. Furthermore, the voice data of the user's speech is collected by a microphone.
[0324] The transmission data processed by the user position detection unit 341, the user direction detection unit 342, and the voice recognition intention understanding unit 343 is transmitted by the communication unit 304 (FIG. 5) to the data management server 10 via the Internet 40. The communication unit 304 also receives response data transmitted from the data management server 10 via the Internet 40.
[0325] The content playback unit 344 plays back the content element based on the received response data. When playing back this content element, not only can the speech (lines) of the character be output from the speaker 332, but also an image of the character can be displayed on the display 331.
[0326] In the data management server 10, the control unit 100 further includes an instruction character selection unit 131, a scenario processing unit 132, and a response generation unit 133. The storage unit 103 (FIG. 3) further stores a character placement DB 161, a position-dependent information DB 162, and a scenario DB 163.
[0327] The communication unit 104 (FIG. 3) receives transmission data transmitted from the playback device 30. The command character selection unit 131 selects a command character by referring to the character placement DB 161 based on the received transmission data, and supplies the selection result to the scenario processing unit 132.
[0328] As shown in FIG. 40, in the character placement DB 161, an arbitrary system and a placement location corresponding to that system are set for each character.
[0329] The scenario processing unit 132 processes the scenario by referring to the position-dependent information DB 162 and the scenario DB 163 based on the selection result from the instruction character selection unit 131 , and supplies the processing result to the response generation unit 133 .
[0330] As shown in Figure 41, the location-dependent information DB162 stores, for each information ID, which is a unique value, its type information, location information such as latitude and longitude, and information about the content linked to the type information and location information.
[0331] As shown in FIG. 42, the scenario DB 163 stores, for each scenario ID that is a unique value, its type information and information related to the content associated with the type information.
[0332] In other words, among the information stored in the character positioning DB 161, the position-dependent information DB 162, and the scenario DB 163, information relating to characters and content corresponds to content elements, system and type information corresponds to context information, and position information corresponds to activation conditions.
[0333] The response generation unit 133 generates response data based on the processing result from the scenario processing unit 132. This response data is transmitted to the playback device 30 via the Internet 40 by the communication unit 104 (FIG. 3).
[0334] In the information processing system 1 configured as described above, the user can set multiple desired voice characters in the scenario, and the user's position and facing direction can be detected in response to an activation condition that indicates a trigger for voice playback, and the voice characters can be switched depending on the detection results.
[0335] Currently, when providing voice character services, when dealing with multiple voice characters, it is difficult to divide up roles among the characters, so as shown in Figure 43, it is necessary to give instructions to each of the voice characters 700A to 700C each time, which is time-consuming.
[0336] On the other hand, in the information processing system 1, when providing a voice character service, the position and direction of the user can be detected and the voice character can be switched depending on the detection result, so that it becomes possible to instruct the voice characters assigned roles to perform desired actions. Therefore, it becomes easy to give instructions to multiple voice characters.
[0337] Specifically, as shown in FIG. 44, the user 900 simply gives instructions to the characters 700A to 700C in the virtual space, and each of the characters 700A to 700C will perform the action in accordance with the instructions given to it.
[0338] 45, user 600 can get an answer to the question from character 700C simply by asking a question by voice in the direction of character 700C in the virtual space. In other words, character 700C can identify information around the position where it is placed, and the user can obtain access rights to the surrounding information due to the presence of character 700C.
[0339] For example, it is possible to realize a user scenario in which voice characters converse with each other, and processing may be added to prevent overlapping conversations through exclusive processing. Furthermore, environmental information surrounding the activation range indicated by the activation condition included in the user scenario may be obtained, and voice may be provided to the user by the voice character specified in that activation range.
[0340] In this way, in the information processing system 1, when the placement of multiple characters can be set, the user can explicitly specify the position of a character in space by specifying the position of the character in the user coordinate system, by specifying the position of the character in the world coordinate system (such as specifying latitude and longitude or a landmark), or by specifying the position of the character within a device such as a playback device 30 that can display the character.
[0341] For example, by arranging characters in the user coordinate system, it is possible to clarify the character that is the target of an instruction by giving instructions to the character in a space with only sound. Also, for example, by having the user give instructions in the world coordinate system, it is easy to assign roles to each character.
[0342] (Overall processing) Next, with reference to FIG. 46, an overview of information processing in the sixth embodiment will be described.
[0343] The information processing shown in FIG. 46 is realized by at least cooperation between the data management server 10 (control unit 100 thereof) and the playback device 30 (control unit 300 thereof) in the information processing system 1.
[0344] 46, in the information processing system 1, sensor data is acquired by real-time sensing (S601). It is determined whether the information obtained from this sensor data satisfies the conditions for invoking a user scenario stored in the user scenario DB 153 (S602).
[0345] If it is determined in the determination process of step S602 that the activation condition is satisfied, it is further determined whether or not only one condition is satisfied (S603).
[0346] If it is determined in the determination process of step S603 that there is only one condition, a content element corresponding to the context information that satisfies the activation condition is presented (S604).
[0347] Furthermore, if the judgment process in step S603 determines that there are multiple conditions, the rules that determine the order in which content elements are presented are referenced (S605), and content elements corresponding to the context information that satisfies the relevant activation conditions are presented in accordance with those rules (S604).
[0348] As this rule, the order of content elements to be presented can be determined from a plurality of content elements according to the orientation of the user estimated from the sensor data (S611, S605).
[0349] Also, as shown in Fig. 38, only content elements in a specific direction may be presented according to the user's direction estimated by sensor data (S621). Furthermore, as shown in Fig. 35, only content elements set in a specific position may be presented according to the user's position estimated by sensor data (S631).
[0350] For example, when the user is facing in a first direction, a content element corresponding to a first character can be identified and presented to the user, and when the user is facing in a second direction, a content element corresponding to a second character can be identified and presented to the user.
[0351] The sixth embodiment has been described above.
[0352] 7. Seventh Embodiment
[0353] The content element playback device 30 may be a single device or a plurality of devices operating in conjunction with each other.
[0354] A case where the playback device 30 is a single device may be, for example, a case where sound is played back from stereo earphones worn by a user outdoors.
[0355] In this case, if the environmental sounds around the user can be superimposed on the content elements and presented simultaneously, it is possible to further enhance the sense of consistency and integration between the provided content elements and the real world around the user. Means for providing the environmental sounds around the user include, for example, open-type earphones that can transmit ambient sounds directly to the ears, and closed-type earphones that capture environmental sounds using a sound collection function such as a microphone and superimpose them as audio data.
[0356] In addition, to ensure consistency with the sense of approach and departure that accompanies user movement, such as walking, it is possible to present the effect of gradually increasing or decreasing the volume (fade in or fade out) when content elements start or stop playing.
[0357] On the other hand, a case in which a plurality of devices including the playback device 30 cooperate to present a content element may occur, for example, when at least one content element is played back by a plurality of devices arranged in an indoor facility.
[0358] At this time, there are cases where one device is assigned to one content element and cases where multiple devices are assigned to one content element.
[0359] For example, by placing three speakers around the user, one for the character's lines, one for the chatter of the cafe, and one for background music, a three-dimensional acoustic environment can be presented.
[0360] The lines of the voice character (FIG. 45, etc.) in the sixth embodiment described above can also be played back from earphones worn by the user. In this case, if the earphones are open type, the user can simultaneously hear sounds from other speakers around the user, making it possible to present linked content elements.
[0361] Furthermore, the voice of the voice character may be localized to a specific position, and the appearance of the voice character may be displayed on a display in the vicinity of the position. This appearance display service may be provided as a paid service.
[0362] Alternatively, character A's lines can be played by detecting the speaker installed in the closest position out of three speakers, and can be made to follow the user's movements so that they are played from the closest speaker.
[0363] To enable such operations, the device has a means for grasping its own location relative to the user's location or the location of other devices. One example of such a means is to install a camera with a function capable of transmitting a blinking code of an LED (Light Emitting Diode) to each pixel installed indoors, and by providing each playback device with a coded light emission transmission function using at least one or more LEDs, it is possible to simultaneously obtain the ID of each device and its assumed location situation.
[0364] Furthermore, functions that can be reproduced by the playback device 30 are registered in advance as device function information in a dedicated database such as a device function information DB, or in the scenario DB 152. Here, a device function describes a playback function that can be realized by a device having one ID, and there are cases where one function is assigned to one device, such as "audio playback" for a speaker, and cases where multiple functions are assigned to one device, such as "image display" and "audio playback" for a television set, or "illumination adjustment" and "audio playback" for a light bulb-type speaker.
[0365] By using this device function information, not only can playback devices in the user's vicinity be identified, but it can also enable a television set to be used as a device for audio playback only, for example. To achieve this, devices that have multiple functions in one device, such as a television set, will have a mechanism that breaks the conventional internal functional coupling within the device and allows each function to function individually and independently based on an external linking signal.
[0366] (Overall processing) Next, with reference to FIG. 47, an overview of information processing in the seventh embodiment will be described.
[0367] The information processing shown in FIG. 47 is realized by at least cooperation between a plurality of devices in the information processing system 1, including the data management server 10 (control unit 100 thereof) and the playback device 30 (control unit 300 thereof).
[0368] As shown in FIG. 47, in the information processing system 1, sensor data is acquired by real-time sensing (S701), and it is determined whether the information obtained from this sensor data satisfies the conditions for invoking the user scenario (S702).
[0369] If it is determined in the determination process of step S702 that the activation condition is met, the process proceeds to step S703. Then, in the information processing system 1, devices capable of presenting the content element are searched for (S703), and at least one device is controlled according to the search result (S704).
[0370] As a result, the content element corresponding to the context information that satisfies the activation condition is presented from one or more devices to be controlled (S705).
[0371] Furthermore, when presenting this content element, the voice of the agent from the content element can be output from headphones worn by the user (electroacoustic transducers worn on the user's ears) (S711), and the appearance of the agent can be displayed on the display (S712).
[0372] In this way, content elements can be presented on one or more devices and through one or more output modalities.
[0373] The seventh embodiment has been described above.
[0374] 8. Eighth Embodiment
[0375] By sharing the scenario currently being used by the user (user scenario) and the contents of the "content element-context information" dataset with external service providers, it is possible to provide services that utilize the content and context that make up the scenario in a collaborative manner.
[0376] As an example, here we will give an example of service collaboration through sharing content elements with restaurants.
[0377] When a user is currently using a scenario that is made up of content elements and context information of a certain animation, the restaurant is provided with information about the contents of the scenario and the fact that the scenario is currently being used.
[0378] In this restaurant, a menu related to the anime, such as omelet rice, is prepared in advance, and a scene is envisioned in which the menu is displayed on an electronic menu that a user using the scenario opens in the restaurant.
[0379] Another example is a service that shares context with an English conversation school.
[0380] As in the previous examples, a scenario can be created in which audio data from an English conversation skit held by an English conversation school is used as the content element, and the situation in which the conversation takes place is set as the context, and the scenario can be provided to the user.
[0381] Furthermore, by sharing only the context information set by the user when using the above-mentioned animation "content element-context information" dataset and providing English conversation skits that fit that context, it becomes possible to provide services at a lower cost. Furthermore, it is possible to design services that expand the points of contact between users, such as by having animation characters read the skits aloud.
[0382] Similarly, it is possible to set up a link between a music streaming service and a restaurant, English conversation school, etc.
[0383] As mentioned above, when a user using a scenario that uses a distributed song or a portion of a song as a content element enters a restaurant, they will be served a drink that matches the restaurant's worldview. Additionally, English conversation skits that fit the context of the song, which does not contain lyrics, are also provided at the same time. Furthermore, it is possible to create and provide new scenarios that combine songs and English conversation, or to have the anime characters used by the user explain the differences between songs or introduce new songs.
[0384] Furthermore, the distribution of context information in the user's daily life space set in a scenario created by another service may be acquired, and music according to the context may be automatically provided as a content element.
[0385] This function allows users to receive music or parts of music that fit the context they have set, for example, on a daily basis, in a location with context information they have set themselves, thereby avoiding the situation of getting bored of listening to the same music every day.
[0386] Furthermore, by obtaining feedback from users such as "likes," the system can constantly acquire information about the relevance of context information and content elements and perform machine learning to improve accuracy.
[0387] (Overall processing) Next, with reference to FIG. 48, an overview of information processing in the eighth embodiment will be described.
[0388] The information processing shown in Figure 48 is realized by at least cooperation between the data management server 10 (control unit 100) and the playback device 30 (control unit 300) in the information processing system 1, as well as servers provided by external services.
[0389] As shown in FIG. 48, in the information processing system 1, at least one content element is extracted from content consisting of multiple media (S801), context information is assigned to each content element, and the content element is stored in the content element-context information DB 151 (S802).
[0390] Then, one or more data sets of "content element-context information" are stored as a scenario in the scenario DB 152 (S803). Also, if a user scenario is generated, it is stored in the user scenario DB 153 (S804).
[0391] The data set, scenario, or user scenario of "content element-context information" accumulated in this way can be provided to external services (S805), which enables providers of external services such as music streaming distribution services to control the services they provide to match the scenario, user scenario, etc. (S811).
[0392] Furthermore, in the information processing system 1, sensor data is acquired by real-time sensing (S821), and it is determined whether or not the information obtained from this sensor data satisfies the conditions for invoking the user scenario (S822).
[0393] If it is determined in the determination process of step S822 that the activation condition is satisfied, a content element corresponding to the context information that satisfies the activation condition is presented (S823).
[0394] At this time, if a scenario, user scenario, etc. is provided to an external service, a service element appropriate for the content element associated with the scenario, user scenario, etc. is selected (S831), and the service element is presented simultaneously with the content element (S832).
[0395] For example, in a music streaming distribution service, a voice character corresponding to a content element (song) associated with a user scenario can be selected (S841), and introduction information can be presented as a DJ introducing the song on the service (S842).
[0396] The eighth embodiment has been described above.
[0397] 9. Ninth Embodiment
[0398] A scenario created by a user (user scenario) can be shared among users using a sharing means.
[0399] Here, social media such as social networking services (SNS) are used as a means of sharing, and scenarios created by users (user scenarios) can be published, for example, on each SNS account, and can be searched and categorized according to the similarity of content elements, similarity of context, similarity of trigger condition settings, etc.
[0400] Here, with regard to the similarity of the settings of the activation conditions, a map application may be used as a means of sharing, allowing the user to discover new scenarios by identifying and presenting scenarios that include the user's current location as an activation condition.
[0401] Information on the works and authors that form the basis of the scenario's content elements, information on the authors who extracted the content elements and added context, and information on the users who set the activation conditions can be linked to the scenario, and users who obtain the scenario can follow their favorite authors and users.
[0402] (Overall processing) Next, with reference to FIG. 49, an overview of information processing in the ninth embodiment will be described.
[0403] The information processing shown in Figure 49 is realized by at least cooperation between the data management server 10 (control unit 100) and the playback device 30 (control unit 300) in the information processing system 1, as well as servers provided by social media.
[0404] As shown in FIG. 49, in the information processing system 1, at least one content element is extracted from content made up of a plurality of media (S901), and context information is assigned to each content element (S902).
[0405] Then, one or more data sets of "content element-context information" are stored as a scenario in the scenario DB 152 (S903). Also, when a user scenario is generated, it is stored in the user scenario DB 153 (S904).
[0406] Scenarios and user scenarios stored in this way can be uploaded to a social media server on the Internet 40 (S905). This allows other users to view the scenarios and user scenarios published on social media (S906). Users can also follow their favorite authors and users regarding the scenarios they have obtained.
[0407] In steps S911 to S913, if the sensor data obtained by real-time sensing satisfies the triggering condition of the user scenario, a content element corresponding to the context information that satisfies the triggering condition is presented.
[0408] The ninth embodiment has been described above.
[0409] <10. Tenth Embodiment>
[0410] In the above-described embodiment, the explanation has focused mainly on audio data and video data, but the data constituting the content elements is not limited to audio and video, but may include formats and data with devices that can present images, tactile sensations, smells, etc., such as playing videos using AR glasses or presenting the tactile sensation of the ground using shoes with a vibration device.
[0411] (Overall processing) Next, with reference to FIG. 50, an overview of information processing in the tenth embodiment will be described.
[0412] The information processing shown in FIG. 50 is executed by the data management server 10 (control unit 100 thereof) in the information processing system 1.
[0413] As shown in FIG. 50, in the information processing system 1, at least one content element is extracted from content consisting of multiple media (S1001), and these multiple media may include at least one of tactile data and smell data that can be presented by the playback device 30.
[0414] The tenth embodiment has been described above.
[0415] <11. Eleventh Embodiment>
[0416] However, since it is possible that the presented content elements may not suit the user, control may be performed to switch the user scenario to another one based on feedback from the user, thereby ensuring that the user is presented with content elements that suit them.
[0417] (Overall processing) An overview of information processing in the eleventh embodiment will be described with reference to FIG.
[0418] The information processing shown in FIG. 51 is realized by at least cooperation between the data management server 10 (control unit 100 thereof) and the playback device 30 (control unit 300 thereof) in the information processing system 1.
[0419] As shown in FIG. 51, in the information processing system 1, at least one content element is extracted from content consisting of a plurality of media (S1101), and context information is assigned to each content element (S1102).
[0420] One or more data sets of "content element-context information" are stored as scenarios in the scenario DB 152. Then, activation conditions are set for the scenarios stored in the scenario DB 152, thereby generating user scenarios (S1103).
[0421] Furthermore, in the information processing system 1, sensor data is acquired by real-time sensing (S1104), and it is determined whether or not the information obtained from this sensor data satisfies the conditions for invoking the user scenario (S1105).
[0422] If it is determined in the determination process of step SS1105 that the activation condition is met, a content element corresponding to the context information that meets the activation condition is presented (S1106).
[0423] Thereafter, if feedback is input from the user (S1107), the user scenario is changed in response to the feedback (S1108). As a result, with the user scenario switched to another one, the above-mentioned steps S1104 to S1106 are repeated, making it possible to present content elements that are more suited to the user.
[0424] Furthermore, by analyzing the feedback input by the user, the user's preferences for content elements are estimated (S1111), and a user scenario is recommended according to the user's preferences (S1121). As a result, the above-mentioned steps S1104 to S1106 are repeated with the recommended user scenario selected, and content elements (e.g., a favorite voice character) that are more suited to the user's preferences can be presented.
[0425] Here, instead of recommending a user scenario, a content element itself may be recommended, and the recommended content element may be presented.
[0426] The eleventh embodiment has been described above.
[0427] <12. Variations>
[0428] In the above description, the information processing system 1 is configured from the data management server 10, the editing device 20, and the playback devices 30-1 to 30-N. However, other configurations may be used, for example, by adding other devices.
[0429] Specifically, the data management server 10 as a single information processing device may be configured as multiple information processing devices by dividing it into a dedicated database server and a distribution server for distributing scenarios, content elements, etc. Similarly, the editing device 20 or the playback device 30 may be configured not only as a single information processing device, but also as multiple information processing devices.
[0430] Furthermore, in the information processing system 1, it is arbitrary which device includes the components (control units) that make up each of the data management server 10, the editing device 20, and the playback device 30. For example, using edge computing technology, part of the information processing by the data management server 10 described above may be executed by the playback device 30, or by an edge server connected to a network close to the playback device 30 (at the periphery of the network).
[0431] In other words, a system refers to a collection of multiple components (devices, modules (parts), etc.), regardless of whether all the components are in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.
[0432] The communication form of each component is also arbitrary. In other words, each component may be connected via the Internet 40 or via a local network (LAN (Local Area Network) or WAN (Wide Area Network)). Furthermore, each component may be connected via a wired connection or wireless connection.
[0433] Conventional technologies aim to simplify use by automating user information search and device operation. This type of automation generally involves determining whether a context classification defined by the system matches a context inferred by sensing the user's behavior and state.
[0434] Such a system is composed of the elements shown in (a) to (d) below, and is characterized by identifying a context defined by the system based on the results of sensing the user's behavior, operations, and physical state.
[0435] (a) Directly analyze and recognize context from sensing data of user behavior (b) Recognizing the content accessed by the user and recognizing the context by analyzing the attribute data and content of that content (c) Have a database of context and content combinations (d) It is assumed that there is a database that associates sensing data with context.
[0436] However, with conventional technology, if a user's behavioral objectives are fixed within a service and tasks and operations are based on certain rules, the user's context can be defined on the system side, making it easier for the user to agree to the context defined by the system.
[0437] On the other hand, when content is presented in a distributed and coordinated manner that adapts to the user's daily life, the user's context is diverse and each user's unique environment changes dynamically, making it difficult for the user to accept the context defined by the system.
[0438] However, the sense of congruence with the context that users feel is subjective and evolving, and it is extremely difficult to predict and adapt this through objective and statistical processing of ex-post data on the context definition defined by the system. Even if this were possible, the accumulation of a huge amount of data would be necessary, and the investment required before the service launch would be unrealistic.
[0439] Furthermore, content presented using conventional technologies is presented to users without changing the format used in conventional services. For example, data or music selected and provided based on context recognition is presented to users in the same format as when it was delivered to the service, without changing the format.
[0440] However, when it comes to presenting content to users in their daily lives, the above-mentioned delivery formats are designed based on traditional viewing behavior, which can hinder free and diverse user behavior in daily life. For example, content such as movies and music is in a format that requires viewers to sit in front of a screen or speakers to watch it, and if it is designed based on traditional viewing behavior, it could hinder user behavior.
[0441] Furthermore, because conventional devices are designed based on traditional viewing behavior, each device is optimized to provide a specific service, and the current situation is that these conventional devices often lack a mechanism for cooperating and sharing some of their functions to adapt to the user's daily behavior.
[0442] For example, while mobile devices such as smartphones have been adapted to users' daily activities through their portability, the premise of screen-centered viewing behavior remains unchanged. For example, when walking on public roads or in public facilities, the so-called "smartphone walking" is considered dangerous due to its characteristic of depriving people of their sight and hearing.
[0443] The above-mentioned Patent Document 1 discloses a device that estimates landmarks that a user is viewing and uses that information to provide a navigation service that indicates the user's direction of travel. However, it does not disclose or suggest the ability to set activation conditions for each user for a context, as in the present technology.
[0444] Furthermore, Patent Document 2 discloses a system that extracts context information and content information from content items, generates an index, and generates recommendations in response based on the user's context and the content of the user's query. However, in Patent Document 2, the context information includes searches, recently accessed documents, running applications, and activity times, but does not include the user's physical location (see paragraph
[0011] ).
[0445] Furthermore, Patent Document 3 discloses a processing device that automatically performs editing when content contains multiple human faces as multiple objects (including audio), by enlarging only the faces of two people defined as context information to a specified size. However, it does not disclose or suggest the technology of the present invention, which records context and audio in association with each other based on content and reuses the recorded content.
[0446] Furthermore, Patent Document 4 discloses that by learning in advance the correspondence between the viewer context (time period, day of the week, etc.) suitable for viewing content and the feature amount of the content based on the content broadcast schedule and broadcast history information and generating a correspondence table of "context-content feature amount," information indicating the context suitable for viewing new content is generated and attached as metadata. However, Patent Document 4 does not disclose how to extract content from existing content.
[0447] Furthermore, in Patent Document 5, context information extracted from sensing data indicating the user's state (motion, voice, heart rate, emotion, etc.) and the video the user is watching at the time are all recorded, and content corresponding to the user's state is extracted using the context information indicating the current user's state, and by generating context information indicating that "the user raised his arms in excitement while watching a soccer broadcast," previously recorded content can be extracted and provided to the user according to keywords such as soccer and excitement, heart rate, and arm motion. However, Patent Document 5 does not disclose extracting content and context from existing content.
[0448] As such, even if the technologies disclosed in Patent Documents 1 to 5 are used, it is difficult to say that a good user experience can be provided when providing services using context information, and there is a demand for providing a better user experience.
[0449] Therefore, this technology uses context information to provide services, allowing users living in different locations to each use a single scenario, thereby providing a better user experience.
[0450] <13. Computer Configuration>
[0451] The above-described series of processes (information processing in each embodiment, such as the information processing in the first embodiment shown in FIG. 6) can be executed by hardware or software. When the series of processes is executed by software, a program constituting the software is installed in a computer of each device. FIG. 52 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0452] In the computer, a CPU (Central Processing Unit) 1001, a ROM (Read Only Memory) 1002, and a RAM (Random Access Memory) 1003 are interconnected by a bus 1004. An input / output interface 1005 is further connected to the bus 1004. An input unit 1006, an output unit 1007, a recording unit 1008, a communication unit 1009, and a drive 1010 are connected to the input / output interface 1005.
[0453] The input unit 1006 includes a microphone, keyboard, mouse, etc. The output unit 1007 includes a speaker, display, etc. The recording unit 1008 includes a hard disk, non-volatile memory, etc. The communication unit 1009 includes a network interface, etc. The drive 1010 drives a removable recording medium 1011 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.
[0454] In a computer configured as described above, the CPU 1001 loads a program recorded in the ROM 1002 or the recording unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executes it, thereby performing the above-mentioned series of processes.
[0455] The program executed by the computer (CPU 1001) can be provided by being recorded on a removable recording medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0456] In a computer, the program can be installed in the recording unit 1008 via the input / output interface 1005 by inserting the removable recording medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the recording unit 1008. Alternatively, the program can be installed in the ROM 1002 or the recording unit 1008 in advance.
[0457] Here, in this specification, the processing performed by a computer according to a program does not necessarily have to be performed chronologically in the order described in the flowchart. In other words, the processing performed by a computer according to a program also includes processing executed in parallel or individually (for example, parallel processing or object-based processing). Furthermore, the program may be processed by one computer (processor), or may be processed in a distributed manner by multiple computers.
[0458] It should be noted that the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the present technology.
[0459] In addition, each step of the information processing in each embodiment can be executed by one device or can be shared and executed by multiple devices. Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0460] The present technology can be configured as follows.
[0461] (1) Context information is pre-mapped to content elements, For each user, it is possible to set an activation condition for at least the context information, and to generate a user scenario consisting of a data set of the context information and the activation condition; When sensor data obtained by sensing a user in real time satisfies an activation condition set in the user scenario, a content element associated with context information corresponding to the activation condition is controlled to be presented to the user. Equipped with a control unit Information processing system. (2) The control unit From content made up of multiple media, extracting content elements comprising at least some media; generating context information corresponding to the content element based on the content; A correspondence database is generated in which the content elements and the context information are stored in association with each other. The information processing system according to (1) above. (3) The control unit generates a scenario database in which a data set consisting of the content elements and the context information is packaged and accumulated based on a certain theme. The information processing system according to (2) above. (4) the content element is part of a streaming content; Information indicating the ID and playback range of the content is stored in association with the context information. The information processing system according to (2) above. (5) The control unit presents another content element including a specific voice character corresponding to the context information before playing the content element. The information processing system according to (4) above. (6) The control unit assigns content information to new content elements by performing machine learning on the relationship between the content elements stored in the correspondence database and the context information. The information processing system according to any one of (2) to (5). (7) The control unit presenting a scenario consisting of a dataset of said content elements and said context information together with map information; An interface is presented that allows a creator who creates a scenario to set a predetermined area on a map as a default value for the activation condition corresponding to the context information. The information processing system according to (3) above. (8) The control unit From the content of the first media, generating a second media different from the first media as a content element; generating context information corresponding to the content element based on the content; A correspondence database is generated in which the content elements and the context information are stored in association with each other. The information processing system according to any one of (1) to (7). (9) the first media includes text; The second media includes TTS (Text To Speech) audio. The information processing system according to (8) above. (10) The control unit A relationship between the first media and the second media is machine-learned in advance, and the second media is generated from the first media based on the results of the machine learning. The information processing system according to (8) or (9). (11) The control unit With respect to the context information, Currently, it is possible to set trigger conditions according to sensor data obtained by sensing the user, and generate a user scenario database consisting of multiple data sets of the context information and the trigger conditions. The information processing system according to any one of (1) to (10). (12) The control unit sets an activation condition according to the captured image data. The information processing system according to (11) above. (13) The control unit sets an activation condition according to the sensor data at that time in response to the characteristic operation of the user. The information processing system according to (11) above. (14) The control unit machine learning the relationship between the context information and the activation condition; Output information according to the results of the machine learning The information processing system according to any one of (11) to (13). (15) The control unit generates context information for a specific activation condition according to the result of the machine learning. The information processing system according to (14) above. (16) The control unit sets an activation condition corresponding to a user for specific context information according to the result of the machine learning. The information processing system according to (14) above. (17) In the sensing, data on which a temporal or spatial activation condition or an activation condition according to a user's behavior can be set is acquired as the sensor data. The information processing system according to any one of (11) to (16). (18) The control unit presenting a scenario consisting of a data set of the content elements and the context information that are pre-associated with map information; An interface is presented that allows the user to set a predetermined area on a map as an activation condition corresponding to the context information. The information processing system according to any one of (1) and (11) to (17). (19) When the same activation condition is set in a plurality of pieces of context information, the control unit presents to the user a plurality of content elements corresponding to the plurality of pieces of context information in accordance with a predetermined rule. The information processing system according to any one of (1) to (18). (20) The control unit identifies one content element from the plurality of content elements according to the orientation of the user estimated from the sensor data, and presents the identified content element to the user. The information processing system according to (19) above. (twenty one) The control unit When the orientation of the user estimated from the sensor data is a first orientation, identifying a content element corresponding to a first character and presenting the content element to the user; When the user's orientation is in a second orientation, a content element corresponding to the second character is identified and presented to the user. The information processing system according to (20) above. (twenty two) The control unit provides information associated with a location of the first character or the second character according to the location of the first character or the second character. The information processing system according to (21) above. (twenty three) The control unit When the sensor data satisfies the activation condition, searching for a device around the user's current location that can present a content element associated with context information according to the activation condition; Controlling the device so that the content element is presented to the user The information processing system according to any one of (1) to (22). (twenty four) The control unit controlling an electroacoustic transducer attached to the ear of the user so that the voice of the agent included in the content element is presented to the user; Controlling a display disposed around a user so that the appearance of the agent included in the content element is presented to the user. The information processing system according to (23) above. (twenty five) The control unit provides a specific user scenario to a service provider via a communication unit. The information processing system according to any one of (1) to (24). (26) The control unit provides the specific user scenario to a music streaming distribution service provider via a communication unit, and sets a voice character corresponding to a content element associated with the specific user scenario as a disc jockey (DJ) who introduces music in the music streaming distribution service. The information processing system according to (25) above. (27) The control unit uploads the user scenario to social media via the communication unit, enabling the scenario to be shared with other users. The information processing system according to any one of (1) to (24). (28) The content elements include at least one of tactile data and smell data that can be presented by a device. The information processing system according to any one of (1) to (27). (29) The control unit switches the user scenario to another user scenario in response to feedback from a user to whom the content element is presented. The information processing system according to any one of (1) to (28). (30) The control unit analyzes the feedback to estimate the user's preferences for the content elements. The information processing system according to (29) above. (31) The control unit recommends the content element or the user scenario according to the user's preference. The information processing system according to (30). (32) The information processing device Context information is pre-mapped to content elements, For each user, it is possible to set an activation condition for at least the context information, and to generate a user scenario consisting of a data set of the context information and the activation condition; When sensor data obtained by sensing a user in real time satisfies an activation condition set in the user scenario, a content element associated with context information corresponding to the activation condition is controlled to be presented to the user. Information processing methods. (33) Computer, Context information is pre-mapped to content elements, For each user, it is possible to set an activation condition for at least the context information, and to generate a user scenario consisting of a data set of the context information and the activation condition; When sensor data obtained by sensing a user in real time satisfies an activation condition set in the user scenario, a control unit controls so that a content element associated with context information corresponding to the activation condition is presented to the user. A computer-readable recording medium that records a program for functioning. [Explanation of symbols]
[0462] 1 Information processing system, 10 Data management server, 20 Editing device, 30, 30-1 to 30-N Playback device, 40 Internet, 100 Control unit, 101 Input unit, 102 Output unit, 103 Memory unit, 104 Communication unit, 111 Data management unit, 112 Data processing unit, 113 Communication control unit, 131 Presentation character selection unit, 132 Scenario processing unit, 133 Response generation unit, 151 Content element-context information DB, 152 Scenario DB, 153 User scenario DB, 161 Character placement DB, 162 Position-dependent information DB, 163 Scenario DB, 200 Control unit, 201 Input unit, 202 Output unit, 203 Memory unit, 204 Communication unit, 211 Editing processing unit, 212 Presentation control unit, 213 Communication control unit, 221 mouse, 222 keyboard, 231 display, 232 speaker, 300 control unit, 301 input unit, 302 output unit, 303 memory unit, 304 communication unit, 305 sensor unit, 306 camera unit, 307 output terminal, 308 power supply unit, 311 playback processing unit, 312 presentation control unit, 313 communication control unit, 321 button, 322 touch panel, 331 display, 332 speaker, 341 user position detection unit, 342 user direction detection unit, 343 voice recognition intention understanding unit, 344 content playback unit, 1001 CPU
Claims
1. Computer, an interface unit that displays a geofence area corresponding to a first activation condition, which is a condition for playing a sound content element among a plurality of sound content elements, and receives input; modifying the first activation condition based on an input of a change to the geofence area; a control unit configured to store, in a storage unit, an activation condition including the changed first activation condition for the sound content element based on the change of the first activation condition; Equipped with the activation condition is set so as to reproduce the audio content element when sensor data from a sensing unit satisfies the activation condition; An information processing program for functioning as an information processing device.
2. An information processing program as described in claim 1, which causes the control unit to function so as to associate each of the activation conditions, including the first activation condition, with each of the multiple audio content elements that make up the audio content.
3. 2. The information processing program according to claim 1, for causing the activation condition to function as a spatial activation condition, a temporal activation condition, an activation condition according to surrounding environmental information, or an activation condition according to user behavior.
4. The control unit sets the activation condition based on a user input, 4. The information processing program according to claim 3, for causing the audio content element to function so as to be reproduced when the spatial activation condition and the temporal activation condition are satisfied.
5. The information processing program according to claim 1 , for causing the program to function such that the change to the geofence area is a change to at least one of the position, size, and shape of the geofence area.
6. the first activation condition is a spatial activation condition, 4. The information processing program according to claim 3, wherein the interface unit is caused to function to accept input of a predetermined area on a map via the interface.
7. The information processing program of claim 3, for causing the control unit to function as follows: to acquire input of a second activation condition that is different from the first activation condition; and to associate the activation condition with context information and store it based on the input of the first activation condition and the second activation condition.
8. The control unit Extracting content elements that are at least some of the media from content consisting of multiple media; Obtaining context information corresponding to the content element based on the content; storing the content element and the context information in association with each other; The information processing program according to claim 1 , for causing the information processing program to function so as to associate the context information with the activation condition and store the association therebetween.
9. the content element is part of the broadcast content; 9. The information processing program according to claim 8, for causing the control unit to function so as to store information indicating a content ID and a playback range corresponding to the content element in association with the context information.
10. Further comprising an acquisition unit for acquiring audio content, The information processing program according to claim 1, for causing the interface unit to function as follows: displaying candidates for the audio content elements corresponding to the acquired audio content via an interface; and accepting input regarding the audio content elements.
11. the interface unit displays candidates for context information via an interface and receives input of selected context information; The information processing program according to claim 8 , for causing the control unit to function so as to associate the context information with the content element and store the associated information based on input of the context information.
12. The information processing program according to claim 11, for causing the control unit to function to generate a scenario database in which a data set consisting of the content elements and the context information is packaged and accumulated based on a certain theme.
13. The information processing program of claim 12, for causing the control unit to function as follows: to present context information for content elements by machine learning the relationship between content elements stored in the scenario database and the context information.
14. The information processing program of claim 1, which causes the control unit to function to acquire input of playback conditions, which are at least one of audio content elements, playback range, volume, repeat playback, fade-in / out, and playback priority level, via the interface unit, and save them together with the activation conditions.
15. The information processing device Displaying a geofence area corresponding to a first activation condition, which is a condition for playing an audio content element among audio content elements consisting of a plurality of audio content elements, and receiving an input; changing the first activation condition based on an input of a change to the geofence area; storing an activation condition including the changed first activation condition for the sound contents element in a storage unit based on the change of the first activation condition; Including, the activation condition is set so as to reproduce the audio content element when sensor data from a sensing unit satisfies the activation condition; Information processing methods.
16. An interface unit that displays a geofence area corresponding to a first activation condition, which is a condition for playing a sound content element among a plurality of sound content elements, and accepts input; modifying the first activation condition based on an input of a change to the geofence area; a control unit configured to store, in a storage unit, an activation condition including the changed first activation condition for the sound content element based on the change of the first activation condition; Equipped with the activation condition is set so as to reproduce the audio content element when sensor data from a sensing unit satisfies the activation condition; Information processing system.
17. The control unit associates the activation conditions including the first activation condition with each of a plurality of audio content elements that make up the audio content.
17. The information processing system according to claim 16.
18. an acquisition unit that acquires sensor data, When the sensor data satisfies the activation condition, the control unit controls output of an audio content element to an output unit according to the activation condition.
17. The information processing system according to claim 16.
19. a transmission unit that transmits a scenario including the audio content elements, context information, and the activation conditions to a server; the server storing the scenario and transmitting the scenario to a device capable of reproducing the audio content element in response to a request from the device; The information processing system according to claim 16, further comprising:
20. The change to the geofence area is a change to at least one of the position, size, and shape of the geofence area.
17. The information processing system according to claim 16.
Citation Information
Patent Citations
Manufacture of optically active secondary aryl amine
JP1989063529A
Information processor, information processing method and program
JP2007172524A
Device, system, and method for context management
JP2015210818A
Multi-household support
JP2017501461A
Information processing device, information processing method, and program
JP2018106444A