Information processing device and information processing method

WO2026196417A1PCT designated stage Publication Date: 2026-09-24NTT DOCOMO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/010386
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2026-09-24

Smart Images

  • Figure JP2025010386_24092026_PF_FP_ABST
    Figure JP2025010386_24092026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to an embodiment includes: an acquisition unit that acquires position information of a user to whom an edited moving image content will be provided; and an output unit that outputs instruction information for instructing a moving image editing device to edit a moving image content, the instruction information corresponding to the position information and including a reproduction time.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing apparatus and information processing method

[0001] The present invention relates to a technology for editing video content.

[0002] For example, Patent Document 1 discloses a video editing apparatus that edits a video to have a length that a user wants to watch during spare time such as travel time or break time and include a portion that the user wants to view.

[0003] Japanese Patent No. 7390877

[0004] When editing a video, the apparatus of Patent Document 1 merely allows the user to specify the desired video length (reproduction time, degree of summary), and cannot perform editing that conforms to the user's situation and constraints, such as the spare time available for the user to watch the video.

[0005] The present invention provides a technology for editing video content according to a user's situation.

[0006] An information processing apparatus according to an aspect of the present disclosure includes: an acquisition unit configured to acquire position information of a user to whom edited video content is to be provided; and an output unit configured to output instruction information for instructing a video editing apparatus to edit the video content, the instruction information including a reproduction time corresponding to the position information.

[0007] An information processing method according to another aspect of the present disclosure includes the steps of: acquiring position information of a user to whom edited video content is to be provided; and outputting instruction information for instructing a video editing apparatus to edit the video content, the instruction information including a reproduction time corresponding to the position information.

[0008] According to the present invention, video content can be automatically edited according to a user's spare time and the like.

[0009] A diagram illustrating the overview of an information processing system. A diagram illustrating the system overview of an information processing system. A diagram illustrating the functional configuration of an information processing device. A diagram illustrating the hardware configuration of an information processing device. A flowchart illustrating the operation in a video distribution system. A diagram illustrating the results of determining the mode of transport and estimating the viewing time and concentration level. A diagram illustrating prompts based on viewing time and concentration level. A diagram illustrating the home screen. A diagram illustrating thumbnails displayed on the home screen.

[0010] 1. Overview Figure 1 is a diagram illustrating the overview of an information processing system. In this example, the information processing system 100 is a system that distributes video content, for example, via a network. The video content includes, for example, movies, dramas, animations, sports, variety shows, etc. The user is a user who watches any video content on the information processing system 100. In this example, the information processing system 100 incorporates a function for editing video content. The information processing system 100 can obtain summary content by editing video content. Summary content is content that summarizes the video content into a short time; in short, it should be generated based on the original video content and have a shorter playback time than the original video content. That is, it is not limited to a method of extracting and reconstructing some of the components of the original video content, but may also include elements that were not present in the original video content. The user can watch video content and summary content through the information processing system 100. Editing video content in the information processing system 100 is performed by inputting instruction information into a video editing device. This video editing device is not limited to a physical device but may also be a software system, and one example is an AI that edits videos (hereinafter referred to as "editing AI"). When the editing AI receives video content data and instruction information as input, it outputs edited content data that has been edited according to the instruction information. Instruction information is information that gives editing instructions to the video editing device that is editing the video content, and one example is a so-called prompt. In this example, the instruction information includes instructions for playback time in the summary content.

[0011] Figure 2 is a diagram illustrating the system overview of the information processing system 100. The information processing system 100 includes an information processing device 10 and a user terminal 20. These devices are connected by a network 200. The network 200 is a computer network such as the Internet. The user terminal 20 is a terminal device that plays (or displays) video content. In this example, the user terminal 20 provides the user with a UI (User Interface) for viewing video content. Examples of user terminals 20 include tablet devices, smartphones, and personal computers. The information processing device 10 is a server that accepts content selections from the user via the user terminal 20 and provides the selected content to the user. In this example, the user can select video content and summary content as content. Note that Figure 2 is merely a system overview for solving the problem, and more specific "configuration" and "operation" will be explained in Chapters 2 and 3, respectively.

[0012] 2. Diagram 3 illustrates the functional configuration of the information processing device 10. The information processing device 10 includes a location acquisition unit 11, an output unit 12, an analysis unit 13, a prompt generation unit 14, an input unit 15, and a storage unit 19. The location acquisition unit 11 acquires the user's location information. Location information is information indicating the user's location and may include, for example, latitude and longitude information obtained by the GPS (Global Positioning System) of the user terminal 20, and ticket gate information obtained when passing through a station ticket gate. As location information, for example, information that allows confirmation of the user's movement status is used. In this example, the location acquisition unit 11 inputs the acquired location information into the storage unit 19. The output unit 12 outputs video content and summary content according to the user's selection to the user terminal 20 and provides them to the user.

[0013] The analysis unit 13 analyzes the location information acquired by the location acquisition unit 11. Based on the location information, the analysis unit 13 can, for example, determine the user's mode of transportation and estimate the viewing time and concentration level. Mode of transportation refers to the means the user uses to travel, including, for example, trains, cars, and walking. Viewing time refers to the amount of time the user is able to watch video content, etc. Concentration level refers to the depth of the user's concentration while watching video content, etc. As will be described later, user attribute information (user attributes) may also be used in the analysis of location information by the analysis unit 13.

[0014] The prompt generation unit 14 generates prompts to be input to the editing AI to edit the video content and obtain a summary content, based on the analysis results of the analysis unit 13. A prompt is an example of instruction information that instructs editing of the video content, and includes, for example, information specifying the playback time of the summary content and information specifying the scenes to be included in the summary content. The input unit 15 inputs the prompts generated by the prompt generation unit 14 and the data of the video content to be edited to the editing AI 300 on an external site. The editing AI 300 performs the editing instructed by the prompts according to a trained model and outputs the summary content. In this example, the input unit 15 acquires the output summary content and inputs it to the storage unit 19. In this example, the video content to be edited is video content that is recommended to the user for viewing, for example, one or more video contents that match the user's tastes and preferences. The storage unit 19 stores various data and information necessary for the information processing device 10 to perform processing. The memory unit 19 stores, for example, video content data, user location information, user attribute information, and summary content data.

[0015] Figure 4 illustrates the hardware configuration of the information processing device 10. Physically, the information processing device 10 is configured as a computer including a processor 101, memory 102, storage 103, communication device 104, input device (optional), display device (optional), and a bus connecting these. Each of these devices operates on power supplied from a battery (not shown). In the following description, the term "device" can be read as a circuit, device, unit, etc. The hardware configuration of the information processing device 10 may include one or more of the devices shown in Figure 3, or it may be configured without some of the devices. Alternatively, multiple devices with different enclosures may be connected via communication to constitute the information processing device 10.

[0016] Each function in the information processing device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 101 and memory 102, which allows the processor 101 to perform calculations, control communication by the communication device 104, and control at least one of the reading and writing of data in the memory 102 and storage 103.

[0017] The processor 101 controls the entire computer, for example, by running an operating system. The processor 101 may be composed of a central processing unit (CPU) that includes interfaces with peripheral devices, control devices, arithmetic units, registers, etc. Alternatively, a baseband signal processing unit or a call processing unit may be implemented by the processor 101.

[0018] The processor 101 reads programs (program code), software modules, data, etc., from at least one of the storage 103 and the communication device 104 into the memory 102 and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described later. Functional blocks of the information processing device 10 may be stored in the memory 102 and implemented by control programs that run on the processor 101. Various processes may be executed by one processor 101, or they may be executed simultaneously or sequentially by two or more processors 101. The processor 101 may be implemented by one or more chips. The program may also be transmitted to the information processing device 10 via a telecommunications line.

[0019] Memory 102 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 102 may also be called a register, cache, main memory, etc. Memory 102 can store executable programs (program code), software modules, etc., for carrying out the method according to this embodiment.

[0020] The storage 103 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. The storage 103 may also be called an auxiliary storage device.

[0021] The communication device 104 is hardware (transceiver / receiver device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc.

[0022] Each device, such as the processor 101 and memory 102, is connected by a bus for communicating information. The bus may be configured using a single bus, or different buses may be used for each device.

[0023] The information processing device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by this hardware. For example, the processor 101 may be implemented using at least one of these hardware components.

[0024] In this example, the program stored in the storage 103 includes a program (hereinafter referred to as the "server program") that causes the computer to function as an information processing device 10 in the information processing system 100. When the processor 101 is executing the server program, at least one of the memory 102 and the storage 103 is an example of a storage unit 19, the processor 101 is an example of an analysis unit 13 and a prompt generation unit 14, and the communication device 104 is an example of a position acquisition unit 11 and an output unit 12.

[0025] Although a detailed explanation is omitted, the user terminal 20 is a computer having a processor, memory, storage, communication device, input device, and output device, such as a smartphone, tablet terminal, or personal computer. The user terminal 20 may also have a GPS receiver and a transportation IC card circuit for generating location information. The user terminal 20 has a program installed that causes the computer to function as the user terminal 20 in the information processing system 100 (hereinafter referred to as the "client program").

[0026] 3. Operation Diagram 5 is a flowchart illustrating the operation of the information processing system 100. The operation shown in the flowchart of Figure 5 starts, for example, when a user accesses a video streaming service, or when a specific time arrives (for example, nighttime batch processing). In step S1, the user's location information is transmitted from the user terminal 20 to the information processing device 10, and the information processing device 10 acquires the location information. The user terminal 20 constantly measures the user's location, and the location information transmitted from the user terminal 20 to the information processing device 10 includes information indicating the user's current and past measured locations.

[0027] In step S2, the information processing device 10 analyzes the location information. In this example, the analysis of the location information first involves determining the mode of transportation in step S21. The mode of transportation is determined by the location (or route) indicated by the location information and the speed of travel. Examples of locations (or routes) indicated by the location information include roads, railway tracks, sidewalks, and inside buildings. The mode of transportation may be determined by location alone, but the accuracy of determining the mode of transportation is improved by combining the location indicated by the location information with the speed of travel. For example, if the location information indicates that the user is moving on a road and the average speed of travel is within a specified range (for example, 20 to 40 km / h), the mode of transportation is determined to be "automobile". In addition, ticket gate information included in the location information may be used to determine the mode of transportation, or user attribute information other than the location information (for example, demographic information) may be used, or the current location information and past travel route information may be combined and used.

[0028] After determining the mode of transportation, step S22 estimates the available viewing time and the level of concentration based on the determined mode of transportation. The available viewing time is estimated as the travel time calculated from, for example, the user's past travel records for the determined mode of transportation, multiplied by a coefficient that indicates the proportion of time available for viewing. This coefficient may be set for each user, or a default value may be provided by the system. The level of concentration is estimated by, for example, adjusting a pre-set basic level of concentration associated with the determined mode of transportation according to the travel time. For example, if the travel time is longer than the standard, the level of concentration is adjusted to be relatively higher, and if it is shorter, it is adjusted to be relatively lower. The level of concentration may also be estimated from, for example, past location information and application operation history. For example, a high level of concentration is estimated in locations where there are many operations to access the video streaming service.

[0029] Figure 6 illustrates the results of determining the mode of transportation and estimating the viewing time and concentration level. For example, in the first example, the mode of transportation is determined to be "traveling by car" based on the speed of movement obtained from location information and a comparison of the user's home location and current location included in the user's attribute information. Based on the user's attribute information, such as not having a driver's license, it is determined that the user is sitting in the passenger seat, and the viewing time is estimated to be "30 minutes" with a concentration level of "high". In this example, the seating position within the car is also considered in estimating the concentration level. For example, the concentration level is estimated to be lower in the driver's seat and higher in the passenger seat and back seat. In the second example, the mode of transportation is determined to be "commuting by train" based on ticket gate information included in the location information, and the viewing time is estimated to be "10 minutes" with a concentration level of "low". In the third example, the mode of transportation is determined to be "standing in line while walking around town" based on the speed of movement obtained from location information and the current location, and the viewing time is estimated to be "10 minutes" with a concentration level of "low". In the fourth example, location information and time indicating the user is at home are used to determine that the mode of transportation is "before going to work," the estimated viewing time is "3 minutes," and the level of concentration is estimated to be "high."

[0030] Returning to Figure 5, in step S3, prompts are generated to instruct the editing of the video content based on the estimated viewing time and attention level as described above. In this example, prompts are generated to instruct the playback time based on the viewing time, as well as the editing policy for the video. In this example, the editing policy is determined based on the attention level. Specifically, the types of scenes selected in the editing differ depending on the attention level. For example, if the attention level is at level 1, a prompt is generated to instruct editing to select scenes based on the first policy, and if the attention level is at level 2, which is lower than level 1, a prompt is generated to instruct editing to select scenes based on the second policy. Here, "selecting scenes" means, for example, prioritizing the inclusion of certain types of scenes in the summary. Scenes in the first policy are, for example, scenes that require more attention than scenes in the second policy, and scenes in the second policy are, for example, scenes that are more stimulating than scenes in the first policy. More specifically, scenes in the first policy are, for example, scenes that are important for understanding the story, and scenes in the second policy are, for example, scenes that have visual impact.

[0031] When generating prompts, the user's tastes and preferences may be reflected. For example, if a user is a fan of a particular actor, prompts may be generated that prioritize selecting scenes featuring that actor in movies or other content. Similarly, if a user is a fan of a particular athlete, prompts may be generated that prioritize selecting scenes featuring that athlete in sports programs or other content. Furthermore, if a user is an animal lover, prompts may be generated that prioritize selecting scenes featuring animals. The user's tastes and preferences may be reflected in ways other than scene selection. For example, if a user prefers a leisurely viewing experience, prompts may be generated that suggest longer playback times, while if a user prioritizes efficiency, prompts may be generated that suggest shorter playback times.

[0032] Specifically, prompts are generated by, for example, using a pre-prepared prompt template into which wording tailored to playback time, concentration level, and personal preferences is embedded.

[0033] Figure 7 illustrates prompts based on available viewing time and concentration level. For example, if a user has a high concentration level, a available viewing time of 10 minutes, and a preference for "actor A," a prompt such as "Summarize important scenes, including actor A's lines, within 10 minutes" might be generated. Similarly, if a user has a low concentration level, a available viewing time of 3 minutes, and a preference for "actor B," a prompt such as "Summarize flashy scenes featuring actor B within 3 minutes" might be generated. In other words, important scenes are prioritized when concentration is high, while impactful, flashy scenes are prioritized when concentration is low. If the available viewing time is divided into multiple discrete periods throughout the day, a prompt indicating a playback time based on the sum of these discrete periods may be generated. For example, if past location data confirms that there are three gaps in the available viewing time each day, a prompt indicating a playback time corresponding to the sum of those three gaps might be generated.

[0034] Returning to Figure 5, in step S4, the generated prompt is given to the editing AI to edit the video content and generate summary content. The summary content generated in this way has a length that is appropriate for the time the user can watch and includes scenes that match the user's tastes, preferences and level of concentration. In step S5, information on the home screen, in which the video content and summary content are presented, for example by thumbnails, is transmitted from the information processing device 10 to the user terminal 20, and the home screen is displayed on the user terminal 20. That is, the video content and summary content are presented to the user via the home screen.

[0035] Figure 8 illustrates a home screen. The home screen 1000 is a screen displayed on, for example, the user terminal 20. The home screen 1000 includes a carousel that presents, for example, movie content and variety content by genre. Each carousel displays thumbnails of multiple video contents side by side. In addition, for example, the top carousel on the home screen 1000 displays "recommended" content recommended to the user by the distribution site.

[0036] Figure 9 illustrates a thumbnail displayed on the home screen. In this example, the thumbnail 2000 has a button 2100 for viewing the full video content and a button 2200 for viewing an edited summary of the video content. In this example, the summary content button 2200 is labeled with something like "10-minute summary" to indicate the playback time of the summary content. The playback time of the summary content is adjusted to match the user's available viewing time, thus encouraging the user to watch the summary content.

[0037] Returning to Figure 5, in step S6, the user selects either a video content or a summary content via a thumbnail in one of the carousels on the home screen 1000 on the user terminal 20. Then, selection information indicating the content selected by the user is transmitted from the user terminal 20 to the information processing device 10. In step S7, the information processing device 10, having received the selection information, transmits playback data for the selected content to the user terminal 20, allowing the user to view the content on the user terminal 20. If the user selects a summary content, the summary content viewed is tailored to the user's level of concentration and preferences, making it easier for the user to become interested in the full video content. Furthermore, the playback time of the summary content is adjusted to match the user's available viewing time, allowing the user to finish watching the summary content within their spare time.

[0038] Users who view summary content do not necessarily watch it from beginning to end. For example, if a user watches part of the summary content but loses interest, they may interrupt their viewing and select other content from the home screen 1000. In such cases, the information processing device 10 generates a prompt indicating a playback time corresponding to the remaining time of the summary content, recreates the summary content, and presents it to the user on the home screen 1000. As a result, the user can watch summary content that matches the remaining viewing time after interrupting their viewing.

[0039] 4. Modifications The present invention is not limited to the embodiments described above, and various modifications are possible. Several modifications are described below. Two or more of the matters described below may be applied in combination.

[0040] (1) Analysis Unit The location information analyzed by the analysis unit is not limited to those exemplified in the embodiments described above. Location information may include, for example, stay information regarding places where the user stayed. Stay information may include, for example, information about the stores where the user stayed and information about the duration of stay. The content analyzed by the analysis unit is also not limited to those exemplified in the embodiments described above. The analysis unit may estimate the level of concentration and viewing time directly from the location information without determining the means of transportation. Alternatively, the analysis unit may estimate only the viewing time without estimating the level of concentration.

[0041] The estimation of concentration by the analysis unit is not limited to those exemplified in the embodiments described above. The concentration level may be the concentration level that is pre-set in association with the means of transportation, or it may be estimated independently of the means of transportation, for example, that a high speed of travel corresponds to a low concentration level and a slow speed of travel corresponds to a high concentration level.

[0042] The position information referenced by the analysis unit in analysis may be only the latest position information, or may include past position information. As the past position information, for example, position information of about one week, which is sufficient to confirm the user's past behavior patterns, may be used. The analysis unit may use the latest position information and past position information separately. For example, the analysis unit may use the latest position information for determining the means of transportation, and use the past position information for estimating the concentration level. The concentration level estimated by the analysis unit is not limited to those exemplified in the above-described embodiments. The concentration level may be estimated in three or more levels, for example, or may be estimated as a numerical value from 0% to 100%, for example.

[0043] (2) Editing of video content The editing of video content is not limited to those exemplified in the above-described embodiments. For example, user preferences do not need to be reflected in the editing of video content. The summary content obtained by editing the video content may be, for example, audio content without video, or the presence or absence of video may be switched according to the concentration level. Further, as editing of video content, for example, volume adjustment, sound tone adjustment, BGM selection, or adjustment of color tone (particularly background) may be performed. Instruction information for instructing editing of video content may be in a format other than so-called prompts (text information), for example, an instruction defined by a combination of text information and other information. That is, the AI editing server 300 is not limited to one that handles natural language processing.

[0044] The instruction of playback time in the instruction information is not limited to instructions indicating specific numerical values (such as 10 minutes, 3 minutes), and may be an instruction using an abstract expression. As abstract expressions for instructing the playback time, for example, "watch carefully", "glance through", "relatively long", "relatively short", and the like may be used. The "first policy" and "second policy" for editing video content are not limited to those exemplified in the above-described embodiments. For example, under the first policy, long-take scenes may be selected, and under the second policy, scenes in which a person is captured in a large size may be selected.

[0045] The editing policy for video content is not limited to two, and an editing policy may be determined from three or more editing policies in accordance with factors such as concentration level and user preferences. In addition, the editing policy may be determined by parameters other than concentration level and user preferences. As parameters other than concentration level and user preferences, for example, parameters of user attributes such as age, gender, and occupation may be used, or parameters that vary from day to day such as the day's weather and temperature may be used. User attributes may be generated based on information previously input to and stored in the user terminal 20 by the user, or may be estimated based on the user's past viewing history, and any method for acquiring user attributes may be used.

[0046] The instruction information may be generated according to the analysis result of the position information, or a plurality of pieces of instruction information with different reproduction times and editing policies may be prepared in advance, and the instruction information corresponding to the analysis result of the position information may be input to the video editing apparatus. As a method for presenting the summary content to the user, for example, only the summary content may be presented on the home screen 1000 separately from the main video.

[0047] (3) Part of the functions of the system configuration information processing apparatus 10 (server) may be implemented in the user terminal 20 (client). For example, the function of the analysis unit 13 may be implemented in the client, and the concentration level and viewable time obtained through analysis may be transmitted from the client to the server. Alternatively, the client may perform processing up to the determination of the transportation means and transmit the result to the server, and the server may perform the estimation of the concentration level and the viewable time. In addition, for example, the storage unit 19 may be replaced with a cloud storage service, or the editing AI 300 may be incorporated into the server.

[0048] (4) Others Various programs executed by the processor 101 may be provided by downloading via a network such as the Internet, or may be provided in a state recorded on a computer-readable non-transitory recording medium such as a DVD-ROM. Note that each processor may be, for example, a CPU, an MPU (Micro Processing Unit), or a GPU (Graphics Processing Unit).

[0049] The block diagrams used in the description of the above embodiments show functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A functional block may also be realized by combining software with the one or more of the above devices.

[0050] Functions include, but are not limited to, judgment, decision, judgment, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, assumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning. For example, a functional block (configuration unit) that enables transmission is called a transmitting unit or transmitter. In all cases, as mentioned above, the method of implementation is not particularly limited.

[0051] For example, the information processing device 10 in one embodiment of the present disclosure may function as a computer that performs the processing of the present disclosure.

[0052] Each aspect or embodiment described in this disclosure may be applied to at least one of the following: LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (new Radio), W-CDMA®, GSM®, CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi®), IEEE 802.16 (WiMAX®), IEEE 802.20, UWB (Ultra-WideBand), Bluetooth®, and other appropriate systems, as well as next-generation systems extended based thereon. Furthermore, multiple systems may be applied in combination (for example, a combination of at least one of LTE and LTE-A with 5G).

[0053] The processing procedures, sequences, flowcharts, etc., of each aspect or embodiment described in this disclosure may be reordered, provided they do not contradict each other. For example, the methods described in this disclosure present various step elements in an exemplary order and are not limited to the specific order presented.

[0054] Input and output information may be stored in a specific location (e.g., memory) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be sent to other devices.

[0055] The determination may be made by a value represented by one bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, by comparing with a predetermined value).

[0056] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Therefore, the descriptions in the present disclosure are illustrative and not intended to be restrictive in any way.

[0057] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name. Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technologies (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technologies (such as infrared or microwave), at least one of these wired and wireless technologies is included in the definition of a transmission medium.

[0058] The information, signals, etc., described herein may be represented using any of the following different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc., which may be referred to throughout the above description, may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof. Terms used herein and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meaning.

[0059] Furthermore, the information, parameters, etc., described in this disclosure may be expressed using absolute values, relative values ​​from a predetermined value, or corresponding other information.

[0060] In this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."

[0061] Any reference to elements using the designations “First,” “Second,” etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Accordingly, references to the First and Second elements do not imply that only two elements may be employed, or that the First element must precede the Second element in any way.

[0062] In the above-described configuration of each device, the term "part" may be replaced with "means," "circuit," "device," etc.

[0063] Where the terms “include,” “including,” and variations thereof are used in this disclosure, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to mean exclusive OR.

[0064] In this disclosure, if articles are added by translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.

[0065] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different."

[0066] 100... Information processing system, 10... Information processing device, 11... Position acquisition unit, 12... Output unit, 13... Analysis unit, 14... Prompt generation unit, 15... Input unit, 19... Storage unit, 20... User terminal, 101... Processor, 102... Memory, 103... Storage, 104... Communication device, 1000... Home screen, 2000... Thumbnail, 2100, 2200... Button

Claims

1. An information processing device having an acquisition unit that acquires location information of a user to whom edited video content will be provided, and an output unit that outputs instruction information for instructing a video editing device to edit the video content, the instruction information including playback time corresponding to the location information.

2. The information processing apparatus according to claim 1, further comprising an acquisition unit for acquiring edited video content generated based on the instruction information from the video editing device.

3. The information processing device according to claim 1, wherein the instruction information includes a degree of concentration that the user can concentrate on viewing, estimated based on the location information.

4. The information processing apparatus according to claim 1, wherein the playback time is the sum of the time of a plurality of discrete times.

5. The information processing apparatus according to claim 1, wherein the location information includes past location information, and the playback time is determined based on the past location information.

6. The information processing device according to claim 1, wherein the instruction information is information based on a means of movement determined from the position information.

7. The information processing apparatus according to claim 1, wherein if viewing of edited video content provided to the user is interrupted, it further generates instruction information for editing other video content to produce video content that fits within the remaining playback time.

8. The information processing apparatus according to claim 3, wherein the output unit outputs instruction information to instruct editing to select a scene according to a first policy when the concentration level is at a first level, and outputs instruction information to instruct editing to select a scene according to a second policy when the concentration level is at a second level lower than the first level.

9. The information processing apparatus according to claim 8, wherein the instruction information includes information generated based on the user attributes of the user.

10. An information processing method comprising the steps of: obtaining location information of a user to whom edited video content will be provided; and outputting instruction information for instructing a video editing device to edit the video content, the instruction information including playback time corresponding to the location information.