Video management method, video management device, video management system, and video management program
Patent Information
- Application Number
- JP2025029247
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-09-07
AI Technical Summary
【0009】 本開示によれば、映像データにより容易にコメントを付与できる。
Smart Images

Figure 2026142253000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a video management method, a video management apparatus, a video management system, and a video management program. [Background Art]
[0002] Patent Literature 1 discloses a terminal device including: transmission means for transmitting search condition data indicating search conditions for work content to a server device; reception means for receiving work content data indicating work content transmitted from the server device as a response to the transmission of the search condition data; display instruction means for causing a display device to display the work content indicated by the work content data; instruction acquisition means for acquiring start instruction data indicating a recording start instruction and end instruction data indicating a recording end instruction; and image acquisition means for acquiring image data generated by an imaging device. The transmission means of the terminal device transmits, to the server device, image data generated by the imaging device during a period from the timing according to the start instruction indicated by the start instruction data to the timing according to the end instruction indicated by the end instruction data. [Prior Art Documents] [Patent Documents]
[0003] [Patent Document 1] International Publication No. WO 2016 / 143749 [Summary of the Invention] [Problem to be Solved by the Invention]
[0004] The present disclosure has been devised in view of the above-described conventional situation, and an object thereof is to provide a video management method, a video management apparatus, a video management system, and a video management program that allow easy addition of comments to video data. [Means for Solving the Problem]
[0005] This disclosure provides a video management method performed by a computer equipped with a camera, wherein the computer causes the camera to perform imaging, accepts user input at the timing when comments are to be added during imaging by the camera, generates a tag indicating the timing at which comments are to be added based on the user input, displays the captured video data and the generated tags after imaging by the camera is completed, accepts input operations for comments for each tag, and links the comments with the video data captured by the camera and registers them in a storage medium.
[0006] Furthermore, this disclosure provides a video management device for a system comprising a plurality of video management devices having cameras and a server for managing video data captured by the plurality of video management devices, the video management device comprising a camera, an operation unit capable of receiving user operations, and a display unit for displaying video data captured by the camera, wherein the operation unit receives user operations at the timing when the user wants to add a comment while the camera is capturing video, generates a tag indicating the timing of adding the comment based on the user operation, the display unit displays the captured video data and the generated tag after the camera has finished capturing video, and the operation unit receives a comment input operation for each tag, links the comment with the video data captured by the camera and registers it with the server.
[0007] Furthermore, this disclosure provides a video management system comprising a plurality of terminals having cameras and a server that manages video data captured by the plurality of terminals, wherein the terminals cause the cameras to perform imaging, transmit the captured video data to the server, accept user operations at the timing when the user wants to add comments while the cameras are imaging, generate tags indicating the timing of adding comments based on the user operations and transmit them to the server, the server stores the video data and tags transmitted from the terminals in association, after imaging by the terminals is completed, transmit the captured video data and the tags corresponding to the video data to the terminals, the terminals display the video data and the tags corresponding to the video data, accept comment input operations for each tag, associate the comments with the video data and transmit them to the server, and the server stores the comments with the video data in association.
[0008] Furthermore, this disclosure provides a video management program performed by a processor communicatively connected to a camera, which includes the steps of: causing the camera to perform imaging; receiving user input at the timing when a comment is to be added during imaging by the camera, and generating a tag indicating the timing of the comment addition based on the user input; displaying the captured video data and the generated tag after imaging by the camera has finished; and receiving a comment input operation for each tag, and registering the comment and the video data captured by the camera in a storage medium. [Effects of the Invention]
[0009] According to this disclosure, comments can be easily added to video data. [Brief explanation of the drawing]
[0010] [Figure 1] Block diagram showing an example of the overall configuration of the video management system according to the embodiment. [Figure 2] This figure shows examples of data storage for various types of data in the embodiment. [Figure 3] Sequence diagram showing an example of the imaging procedure for the video management system according to the embodiment. [Figure 4] Sequence diagram showing example comment editing procedure and video search procedure of the video management system according to the embodiment. [Figure 5] A diagram showing an example of the imaging screen and viewing screen. [Figure 6] A diagram showing an example of a viewing screen. [Figure 7] A diagram showing an example of the imaging screen and the AI viewing screen. [Figure 8] A diagram showing an example of a search results screen. [Figure 9] Diagram showing an example of the search results screen and viewing screen. [Figure 10] A diagram showing an example of the image capture screen during voice input. [Figure 11] A diagram showing an example of the viewing screen when using voice input. [Figure 12] Block diagram showing an example of the overall configuration of a video management system according to a modified embodiment. [Figure 13] This figure shows an example of data storage for various types of data in a modified embodiment. [Figure 14] Sequence diagram showing an example of the imaging procedure for a video management system according to a modified embodiment. [Modes for carrying out the invention]
[0011] (Background leading to this disclosure) Traditionally, there is a video sharing technology that involves installing surveillance cameras at each location and sharing the video data captured by these cameras across multiple different locations. In this conventional video sharing technology, when a specific event (e.g., motion detection, intrusion detection, etc.) is detected by the surveillance camera, information indicating this specific event (hereinafter referred to as "tag") is associated with the video data, enabling the searching and viewing of video data capturing that specific event at multiple locations.
[0012] In recent years, with the popularization of mobile terminals equipped with camera and communication functions (for example, smartphones, tablet terminals, etc.), in addition to fixed cameras such as surveillance cameras, video sharing technology using mobile terminals has been demanded. Here, when a mobile terminal is used, unlike a surveillance camera that can automatically generate tags, there has been a problem that an imager must manually perform creation of a tag and input of a comment indicating the content of a specific event.
[0013] For example, Patent Document 1 enables a user to freely specify keywords indicating attributes of work content such as a place name, a worker, or work content in order to realize a keyword search for searching an image of work content that the user desires to view, and stores such tags.
[0014] However, since it is difficult for an imager to input a comment in real time while capturing an image, the imager needs to review the video after capturing to create a tag and input a comment corresponding to the tag, which is very time-consuming. Therefore, the present disclosure describes a video management apparatus, a video management system, and a video management program that can easily add comments to video data.
[0015] Hereinafter, each embodiment that specifically discloses the video management method, video management apparatus, video management system, and video management program according to the present disclosure will be described in detail with reference to the drawings as appropriate. However, an unnecessarily detailed description may be omitted. For example, a detailed description of already well-known matters and a repeated description of substantially the same configuration may be omitted. This is to prevent the following description from being unnecessarily redundant and to facilitate understanding by those skilled in the art. The accompanying drawings and the following description are provided for those skilled in the art to fully understand the present disclosure, and are not intended to limit the claimed subject matter by them.
[0016] (Embodiment) First, the video management system 100 according to the embodiment will be described with reference to Figure 1. Figure 1 is a block diagram showing an example of the overall configuration of the video management system 100 according to the embodiment. Figure 2 is a diagram showing an example of data storage for various types of data in the embodiment.
[0017] The video management system 100 includes user terminals P1,... that accept imaging operations and comment addition operations from users, and a server S1 that manages the association between the images captured by P1,... and comments describing the content of the images on the user terminals. Note that the video management system 100 shown in Figure 1 is just one example and is not limited thereto.
[0018] User terminals P1,... are implemented, for example, by smartphones or tablet devices. User terminals P1,... are connected to server S1 via a network NW, enabling data communication. User terminals P1,... are capable of receiving user operations and, based on those operations, perform image acquisition and comment assignment.
[0019] User terminals P1,... include a communication unit 10, a processor 11, a memory 12, an imaging unit 13, a display unit 14, an operation unit 15, and a sound pickup unit 16.
[0020] The communication unit 10 is connected to the server S1 via a network NW, enabling wireless or wired communication. Wireless communication here refers to communication via a wireless Local Area Network (LAN), such as Wi-Fi (registered trademark). The communication unit 10 outputs various data transmitted from the server S1 to the processor 11. The communication unit 10 also transmits various data output from the processor 11 to the server S1.
[0021] The processor 11 is configured using, for example, a Central Processing Unit (hereinafter referred to as "CPU") or a Field Programmable Gate Array (hereinafter referred to as "FPGA"), and works in cooperation with the memory 12 to perform various processing and control. Specifically, the processor 11 refers to the programs and data held in the memory 12 and executes those programs to realize various functions such as the search unit 111. Note that the functions described above are only a part of the functions that the processor 11 can realize, and are not limited to these.
[0022] For example, when the processor 11 accepts a comment generation or editing operation based on a voice input operation, it performs speech recognition on the user's utterance captured by the operation unit 15 or the sound collection unit 16. The processor 11 may also accept a comment input operation by performing speech recognition on the captured voice data. The method for generating comments using the captured utterance will be described later.
[0023] Alternatively, for example, the processor 11 may use a pre-trained model to detect a predetermined scene from video data, detect that scene, and automatically generate a comment corresponding to the detected scene. In such a case, the processor 11 accepts a user input to select a pre-trained model, and automatically generates a comment by obtaining the pre-trained model specified by the user input from the AI database DB3. The method for generating comments using a pre-trained model will be described later.
[0024] The search unit 111 acquires keywords used for searching for videos based on user input from the operation unit 15. The search unit 111 generates a control command requesting the transmission of video data associated with comments containing the acquired keywords and sends it to the server S1. The search unit 111 acquires the video data transmitted from the server S1 and generates a search results screen SC15A (see Figure 9) showing the search results for videos based on the keywords, and displays it on the display unit 14.
[0025] Memory 12 includes, for example, Random Access Memory (hereinafter referred to as "RAM"), which serves as work memory used when executing various processes of the processor 11, and Read Only Memory (hereinafter referred to as "ROM"), which stores programs and data that define the operation of the processor 11. Data or information generated or acquired by the processor 11 is temporarily stored in RAM. Programs that define the operation of the processor 11 are written to ROM.
[0026] The imaging unit 13 is implemented, for example, by an optical system including a lens and an image sensor. The imaging unit 13 takes images based on control commands output from the processor 11. The imaging unit 13 outputs the captured video data to the processor 11. The video data output here is linked to at least the information of the start time of imaging and information that can identify the imaging unit 13 or the user terminal P1 (hereinafter referred to as "camera ID").
[0027] The display unit 14 is configured using, for example, a Liquid Crystal Display (LCD) or an organic electroluminescence (EL). The display unit 14 displays various screens generated and output by the processor 11 (for example, an image capture screen, a viewing screen, an AI viewing screen, or a search results screen). The display unit 14 may also be integrated with the operation unit 15 and be a user interface configured as a touch panel. In such a case, the display unit 14 accepts user input operations and outputs the results of the user input operations to the processor 11.
[0028] The operation unit 15 is a user interface capable of receiving input operations or voice input operations from the user, and is implemented by, for example, a mouse, keyboard, touch panel, or microphone. The operation unit 15 receives user operations, converts the content of the user operations into electrical signals, and outputs them to the processor 11.
[0029] The sound pickup unit 16 is capable of picking up the user's spoken voice and is implemented by a microphone provided in the user terminal P1,... which also functions as the operation unit 15, or by an external device such as a microphone connected to the user terminal P1,.... The sound pickup unit 16 picks up the spoken voice of the user ordering in real time. The sound pickup unit 16 converts the picked-up voice data into an audio signal (audio data) and outputs it to the processor 11. In addition to the user's spoken voice, the sound pickup unit 16 can also pick up ambient sounds during imaging.
[0030] Server S1 is connected to each of the multiple user terminals P1, ... via a network NW, enabling data communication. Server S1 manages video data captured by each of the multiple user terminals P1, ... and comments associated with the video data. Server S1 includes a communication unit 20, a processor 21, memory 22, a tag information database DB1, and a video storage DB2.
[0031] The communication unit 20 is connected to each of the multiple user terminals P1,... via wireless or wired communication. The communication unit 20 outputs various data transmitted from each of the multiple user terminals P1,... to the processor 21. The communication unit 20 also transmits various data transmitted from the processor 21 to each of the user terminals P1,...
[0032] The processor 21 is configured, for example, using a CPU or FPGA, and works in cooperation with the memory 22 to perform various processes and controls. Specifically, the processor 21 refers to the programs and data held in the memory 22 and executes those programs to realize various functions such as the search unit 211 and the editing and processing unit 212. The functions described above are only a part of the functions that the processor 21 can realize, and are not limited to these.
[0033] The search unit 211 extracts video data from the video storage DB2 that are associated with comments containing the keyword, based on the control commands and keywords transmitted from the user terminals P1,... The search unit 211 transmits the extracted video data to the user terminals P1,...
[0034] The editing and processing unit 212 generates a digest video, a digest video with comments, or a report, etc., based on the specified video data, based on control commands transmitted from user terminals P1,.... In this disclosure, the editing and processing unit 212 is described as a function that can be executed by server S1, for example, but it may also be a function that can be executed by each user terminal P1,.... The digest video, digest video with comments, and report referred to here will be described later.
[0035] Memory 22 includes, for example, RAM as work memory used when executing each process of processor 21, and ROM which stores programs and data that define the operation of processor 21. Data or information generated or acquired by processor 21 is temporarily stored in RAM. Programs that define the operation of processor 21 are written to ROM.
[0036] The tag information database DB1 is a type of storage, consisting of storage media such as flash memory, a hard disk drive (hereinafter referred to as "HDD"), or a solid state drive (hereinafter referred to as "SSD"). The tag information database DB1 stores tag information associated with each video.
[0037] The tag information database DB1 stores, for example, a tag ID, the time the tag ID was generated, the tag type, the content of the comment, a camera ID indicating the user terminal P1 or imaging unit 13 that is the source of the corresponding video data, an AI flag indicating whether or not a trained model is used, and an AI type indicating the type of scene detected by the trained model, all linked together for each tag ID. Note that the tag information database DB1 shown in Figure 2 is just an example and is not limited thereto.
[0038] The tag ID "aaaaaa" shown in Figure 2 indicates that the tag type is "startVideo" and that it is a tag indicating the start of image capture. The video data corresponding to tag ID "aaaaaa" has an image capture start time of "2024 / 1 / 4 10:54:32.123" and the camera ID used for capture is "00001". Furthermore, the video data corresponding to tag ID "aaaaaa" has content (comments) attached that are "video comments", and the AI flag is "0", meaning that a trained model was not used and the content is a manually entered comment.
[0039] Furthermore, the tag ID "bbbbbb" shown in Figure 2 indicates that the tag type is "comment" and that it is a tag indicating the time the comment was added. The video data corresponding to tag ID "bbbbbb" has a comment added at "2024 / 1 / 4 10:57:45.456" and the camera ID that captured the image is "00001". In addition, the video data corresponding to tag ID "bbbbbb" has content (comment) that says "This is a comment", and the AI flag is "0", meaning that a trained model was not used and the content was manually entered as a comment.
[0040] Furthermore, the tag ID "cccccc" shown in Figure 2 indicates that the tag type is "comment" and that it is a tag indicating the time the comment was added. The video data corresponding to tag ID "cccccc" has a comment added at "2024 / 1 / 4 10:57:45.456" and the camera ID that captured the image is "00001". In addition, the video data corresponding to tag ID "cccccc" has an AI flag of "1", meaning that a trained model was used, the content (comment) generated by the trained model is "the floor is wet", and the type of scene (AI type) that the trained model used can detect is "chemical leak".
[0041] The video storage DB2 is a type of storage system, configured using storage media such as flash memory, HDD, or SSD. The video storage DB2 stores video data captured by multiple user terminals P1,...
[0042] The video storage DB2 stores, for example, the TS file name, which is the file name in which the video data is stored, the start time of video data capture, and the camera ID, which indicates the user terminal P1 or imaging unit 13 that is the source of the video data capture. Note that the video storage DB2 shown in Figure 2 is just an example and is not limited thereto. For example, the video storage DB2 may further store information about the storage period for which each video data will continue to be stored in the video storage DB2.
[0043] The TS file name "xxxxxxx.Ts" shown in Figure 2 includes the camera ID "0001" (not shown) on which the video data was captured, the start time of video data capture "2024 / 1 / 4 10:54:32.123" (not shown), and the video data duration "1.003" (not shown) within the TS file name itself. Similarly, the TS file name "yyyyyyy.Ts" shown in Figure 2 includes the camera ID "0001" (not shown) on which the video data was captured, the start time of video data capture "2024 / 1 / 4 10:54:33.126" (not shown), and the video data duration "1.003" (not shown) within the TS file name itself.
[0044] When the processor 21 of server S1 extracts tag information (for example, time information of comment tags, or comment content, etc.) associated with each video data from the tag information database DB1 based on control commands sent from user terminals P1,..., it searches for and extracts tag information associated with the camera ID of user terminals P1,.... Here, each user terminal P1,... captures only one video data that a single camera can capture during the same time period. Based on the camera ID and the capture start time, the processor 21 searches the tag information database DB1 and video storage DB2 for the video data and tag information (comments) that user terminals P1,... request to be transmitted and retrieves them.
[0045] <Example of imaging procedure for a video management system> Next, with reference to Figure 3, an example of the imaging procedure of the video management system 100 in the embodiment will be described. Figure 3 is a sequence diagram showing an example of the imaging procedure of the video management system 100 according to the embodiment. Note that each of the user terminals P1, ... performs the same processing in the example imaging procedure shown in Figure 3. Therefore, in the explanation of Figure 3, an example of the imaging procedure performed by user terminal P1 will be described.
[0046] When the user terminal P1 receives an input from the user to start imaging via the operation unit 15, it starts imaging with the imaging unit 13 (St11). The user terminal P1 generates an imaging start tag that links the imaging start time information with a camera ID that can identify the user terminal P1 or the imaging unit 13, and sends it to the server S1 (St12).
[0047] Server S1 obtains the image capture start tag sent from the user terminal P1. Server S1 associates the image capture start tag with the image capture start time information and the camera ID and stores it in the tag information database DB1 (St13).
[0048] The tag information database DB1 generates a tag ID, which is identification information that makes the tag identifiable (St14). The tag information database DB1 stores the generated tag ID in association with the image acquisition start tag, image acquisition start time, and camera ID output from server S1.
[0049] The video management system 100 repeatedly executes the imaging process LP11 while the user terminal P1 is capturing images. Furthermore, while the user terminal P1 is capturing images, if a user performs a tagging operation to add comments at any time, the video management system 100 executes the tagging process LP12.
[0050] In the imaging process LP11, the user terminal P1 sends video data to the server S1 at predetermined intervals (for example, every second) for storage. The user terminal P1 associates the video data corresponding to the predetermined interval with the information of the start time of the video data acquisition and the camera ID, and sends it to the server S1 (St15). The server S1 generates a file name as identification information that makes the video data output from the user terminal P1 identifiable (St16).
[0051] Server S1 associates the video data, image capture start time, and camera ID transmitted from user terminal P1 with the generated file name and stores them in video storage DB2 (St17). Note that the image capture start time in step St13 is the video recording start time, and the recording start time in steps St15 and St17 is the file recording start time.
[0052] In the tagging process LP12, the user terminal P1 accepts a tagging operation to add comments to the video data being captured, based on user operations using the operation unit 15 (St18). When the user terminal P1 accepts a user operation, it generates a comment tag, links the generated comment tag with comment time information indicating when the user operation was accepted, and sends it to the server S1 (St19).
[0053] Server S1 retrieves the comment tag, comment time information, and camera ID information sent from user terminal P1, and outputs and stores them in the tag information database DB1 (St20).
[0054] When a comment tag is output from server S1, the tag information database DB1 generates a tag ID (St21) that makes this comment tag identifiable. The tag information database DB1 stores the comment tag and comment time information output from server S1, associating it with the generated tag ID.
[0055] When the user terminal P1 receives an operation to end imaging from the user via the operation unit 15, it terminates imaging by the imaging unit 13 (St22). The user terminal P1 generates an imaging termination tag that links the imaging termination time information with a camera ID that can identify the user terminal P1 or the imaging unit 13, and sends it to the server S1 (St23).
[0056] Server S1 receives the image capture completion tag sent from user terminal P1. Server S1 associates the image capture completion tag, the image capture completion time information, and the camera ID and stores them in the tag information database DB1 (St24).
[0057] The tag information database DB1 generates a tag ID, which is identification information that makes the tag identifiable (St25). The tag information database DB1 stores the generated tag ID in association with the image completion tag, image completion time, and camera ID output from server S1.
[0058] As described above, the video management system 100 in this embodiment can accept a tagging operation for the location where a comment is added to the video data, that is, the capture time when the comment is added. This allows the video management system 100 to add a tag that makes it possible to refer to the capture time when a comment is requested after the capture, even if the user who is taking the image cannot input a comment in real time during the capture.
[0059] Next, with reference to Figure 4, an example of the comment editing procedure and the video search procedure of the video management system 100 in the embodiment will be described. Figure 4 is a sequence diagram showing an example of the comment editing procedure and the video search procedure of the video management system 100 according to the embodiment. Note that each of the user terminals P1, ... performs the same processing in the example of the comment editing procedure and the video search procedure shown in Figure 4. Therefore, in the explanation of Figure 4, the example of the comment editing procedure and the video search procedure performed by user terminal P1 will be described.
[0060] <Example of comment editing procedure for video management system> User terminal P1 accepts the selection operation for video data to be edited with comments. User terminal P1 generates a control command requesting the transmission of tag information corresponding to the selected video data and sends it to server S1 (St31).
[0061] Server S1 reads tag information corresponding to the video data from the tag information database DB1 based on the control command sent from the user terminal P1 (St32). Server S1 then sends the read tag information to the user terminal P1 (St33).
[0062] The user terminal P1 generates viewing screens SC12, SC12A, SC17 (see Figures 5, 6, or 11) containing each tag information based on the tag information transmitted from the server S1, and displays them on the display unit 14. The user terminal P1 accepts user operation to select the tagging location that the user requests to tag (St34). The user terminal P1 generates a control command requesting the transmission of video data corresponding to the tagging location (time) selected by the user operation and sends it to the server S1 (St35).
[0063] Server S1 reads video data corresponding to tag information stored in video storage DB2 based on control commands sent from user terminal P1 (St36). Server S1 then sends the read video data to user terminal P1 (St36).
[0064] User terminal P1 acquires video data transmitted from server S1. Based on tag information and video data, user terminal P1 generates viewing screens SC12, SC12A, SC17 (see Figures 5, 6, or 11) that can accept editing operations for comments on the selected video data, and displays them on display unit 14. User terminal P1 accepts comment editing operations by receiving user operations on viewing screens SC12, SC12A, SC17 (St38). As a result of comment editing by the user, user terminal P1 generates a comment tag containing the comment content, associates the generated comment tag with the tag ID, and sends it to server S1 (St39).
[0065] Server S1 outputs and stores the comment tags and tag IDs sent from user terminal P1 in the tag information database DB1 (St40).
[0066] Server S1 stores the comment tags and tag IDs sent from user terminal P1 in the tag information database DB1 and stores or updates comment tags for tag information with the same tag ID (St41).
[0067] As described above, the video management system 100 according to the embodiment can add comments to the video data corresponding to the tag information at the time of imaging, based on the tags generated at the time of imaging.
[0068] <Example of video search in a video management system> User terminal P1 accepts input of keywords to search for the target video data (St51). User terminal P1 generates a control command requesting the transmission of video data with comments containing the entered keywords and sends it to server S1 (St52).
[0069] Server S1 requests the tag information database DB1 to search for comments containing the specified keyword, based on the control command sent from user terminal P1 (St53).
[0070] The tag information database DB1 searches for comments (comment tags) that contain the specified keyword (St54).
[0071] Server S1 reads the tag information associated with comments containing the specified keyword based on the search results (St55) and sends it to user terminal P1 (St56).
[0072] The user terminal P1 generates a search results screen SC15A (see Figure 9) showing search results for comments containing the keyword, based on the tag information transmitted from the server S1, and displays it on the display unit 14. The user terminal P1 accepts a selection operation from among the multiple tag information (e.g., comments or thumbnail images of video data, etc.) included in the search results screen SC15A (see Figure 9) for the video to be viewed by the user (St57). The user terminal P1 generates a control command requesting the transmission of the video data associated with the selected video to be viewed, along with the tag information corresponding to this tag information (e.g., information on the start time of imaging or camera ID, etc.), and sends it to the server S1 (St58).
[0073] Server S1 reads video data corresponding to the tag information of the video being viewed, which is stored in the video storage DB2, based on the control command sent from the user terminal P1. Server S1 then sends the read video data to the user terminal P1 (St59).
[0074] User terminal P1 acquires video data transmitted from server S1. User terminal P1 generates a viewer screen SC15B (see Figure 9) that can play the acquired video data and displays it on display unit 14 (St60). When user terminal P1 receives a video data playback operation from the user using the operation unit 15, it plays the video data (St60).
[0075] As described above, the video management system 100 according to the embodiment can search for video data containing keywords based on comments attached to the video data and display it on the user terminal P1.
[0076] <How to generate comments through user interaction> Next, the method for generating comments through user operation will be explained with reference to Figures 5 and 6, respectively. Figure 5 shows an example of the imaging screen SC11 and the viewing screen SC12. Figure 6 shows an example of the viewing screen SC12A. Note that the imaging screen SC11 and the viewing screens SC12 and SC12A shown in Figures 5 and 6 are examples and are not limited thereto.
[0077] When an application for performing imaging by the imaging unit 13 is launched on the user terminal P1, it generates an imaging screen SC11 and displays it on the display unit 14. The imaging screen SC11 includes the captured image MV11, a comment tag button BT111, a mute button BT112, and an imaging button BT113.
[0078] The captured image MV11 is an image captured by the imaging unit 13. In Figure 5, the area where the captured image MV11 is displayed shows the captured image captured by the imaging unit 13 in real time.
[0079] The comment tag button BT111 accepts a tagging operation, that is, an operation to generate a comment tag. When the comment tag button BT111 is selected (pressed) by the user, the processor 11 generates a comment tag associated with an empty comment. The processor 11 associates the generated comment tag with the time the comment tag button BT111 was selected (pressed) by the user (comment time) and the camera ID, and sends them to the server S1.
[0080] The mute button BT112 accepts the operation of the control unit 15 to stop / resume sound recording during imaging. When the mute button BT112 is selected (pressed) by the user, the processor 11 controls the stopping and resuming of sound recording during imaging by the control unit 15.
[0081] The imaging button BT113 accepts the start / end operation of imaging. When the imaging button BT113 is selected (pressed) by the user, the processor 11 controls the start / end operation of imaging. When the processor 11 receives the start of imaging from the user, it generates an imaging start tag indicating the start of imaging, associates the imaging start time (the timing when the imaging button BT113 was selected (pressed) by the user) with the camera ID, and sends it to the server S1.
[0082] The imaging screen SC11 shown in Figure 5 displays the captured image MV11, which was captured 5 seconds after the start of imaging. When the comment tag button BT111 is selected (pressed) on the imaging screen SC11, the processor 11 generates a comment tag associated with the comment time corresponding to the 5-second mark after the start of imaging. Also, when the capture button BT113 is selected (pressed) on the imaging screen SC11, the processor 11 terminates imaging by the imaging unit 13.
[0083] After the user terminal P1 receives the operation to end imaging from the imaging unit 13, if it receives an operation to specify video data to be edited based on the user's operation, it obtains the specified video data and the tag information already attached to this video data from the server S1. Based on the obtained video data and tag information, the user terminal P1 generates a viewing screen SC12 that can accept the user's comment editing operation and displays it on the display unit 14.
[0084] The viewing screen SC12 includes the captured image MV12, the tag pin PN122, the dropdown SL12, the playback date and time information DT12, the play button BT121, the comment add button BT122, the comment information CM121, CM122, CM123, CM124, the digest video creation button BT123, and the storage period extension button BT124.
[0085] The captured image MV12 is one of several captured images included in the video data, and corresponds to the current playback position PN121 and playback date and time information DT12.
[0086] Tag pin PN122 indicates the time when each comment tag (comment) was assigned to the video data. The multiple tag pins PN122 shown in Figure 5 correspond to the assignment times of comment information CM121 to CM124.
[0087] The dropdown SL12 indicates the duration of one segment of the seek bar, which represents the recording period of the video data and the current playback position. For example, if the selected value in dropdown SL12 changes from "1 minute" to "10 minutes," one segment of the seek bar changes from 1-minute increments (e.g., "10:12" to "10:13" as shown in Figure 5) to 10-minute increments (e.g., "10:12" to "10:22"). In addition, the viewing screen SC12 changes the size of the bar representing the recording period of the video data (the dot hatch area displayed between the four tag pins as shown in Figure 5), the display position of the current playback position PN121, and the tag pin PN122 in accordance with the change in the value selected in dropdown SL12.
[0088] For example, the viewing screen SC12 shown in Figure 5 visualizes the tagged comment information CM121-CM124 for the one minute from "10:12" to "10:13", which includes the current playback position (time) "10:12:45", as well as the time each of these comment information CM121-CM124 was assigned.
[0089] The playback button BT121 accepts user input for playback / stop of video data. When the playback button BT121 is selected (pressed), the processor 11 controls the playback and stopping of the captured image (video data) displayed in the captured image MV12. The playback button BT121 may also allow the user to select playback methods such as double-speed playback or time-lapse playback. This allows the user to more easily recall the content of the video data through time-lapse playback.
[0090] Furthermore, when playing back at double speed or in time-lapse mode, the processor 11 may play back only a predetermined period of time (for example, 3 seconds or 5 seconds) before and after the time the comment was added at normal speed (i.e., 1x speed). This allows the processor 11 to reduce the chances of the user missing the content of the added comment and the video at the time the comment was added while viewing the video.
[0091] The comment addition button BT122 accepts the operation to add a comment tag (comment). When the comment addition button BT122 is selected (pressed) by the user, the processor 11 generates a comment tag at the position (time) corresponding to the current playback position PN121. Note that the operation to add a comment tag (comment) is not limited to the selection (pressing) of the comment addition button BT122. The operation to add a comment tag (comment) may be achieved by a specific operation, such as double-tapping the area on the display unit 14 where the captured image is displayed.
[0092] Each of the comment information entries CM121 to CM124 contains information about a comment tagged in the current video data display section. In the example shown in Figure 5, the comment information entries CM121 to CM124 include information about the location (time) where the comment tag was assigned and the comment content. Specifically, comment information CM121 indicates that the comment was assigned to the video data at the time "10:12:45" and that the comment content is "Sample video!". When the processor 11 receives a user operation on the area where the comment content is displayed, it accepts an editing operation for the comment content.
[0093] The digest video creation button BT123 accepts an operation to generate a digest video by cutting out predetermined sections before and after the position (time) where each comment tag is attached, and then connecting the cut-out predetermined sections. When the digest video creation button BT123 is selected (pressed) by the user, the processor 11 generates a control command requesting the generation of a digest video of this video data and sends it to the server S1.
[0094] The editing and processing unit 212 of server S1 extracts a predetermined section (for example, a few seconds before and after) that includes the time when a comment tag was assigned, based on control commands sent from user terminals P1,... The editing and processing unit 212 generates a digest video that includes comments (i.e., text data) corresponding to the extracted video data, and connects the extracted video data in chronological order of the comment tags.
[0095] The method for generating the digest video is not limited to the examples described above. The editing and processing unit 212 may extract and combine video data to a predetermined length (e.g., 3 minutes, 5 minutes, etc.) based on the time when the comment tag was added, and generate a digest video that includes the corresponding comments (text data). Alternatively, the editing and processing unit 212 may simply extract the captured images corresponding to the timing when the comment tag (comment) was added, use the extracted captured images as thumbnail images, and generate a report containing the comments corresponding to each thumbnail image as a digest video.
[0096] The generated digest video may be sent to user terminals P1, ... and downloaded, or stored in the video storage DB2. Alternatively, after storing the generated digest video in the video storage DB2, the editing and processing unit 212 may generate and output a 2D barcode that allows viewing of the digest video. This allows the editing and processing unit 212, when creating a report on inspection work, to attach a digest video containing the video data of the inspection work and comments attached to the video data as a 2D code to the report. In this way, the editing and processing unit 212 can support the creation of reports with attached videos by generating a digest video and a 2D code that allows viewing of the digest video. Furthermore, the editing and processing unit 212 may output the comments as the report body (text data), link this text body with the digest video, and generate and output the report data.
[0097] The storage period extension button BT124 accepts an operation to extend the storage period for video data displayed on the viewing screen SC12 before it is stored in the video storage DB2. When the storage period extension button BT124 is selected (pressed) by the user, the processor 11 generates a control command requesting an extension of the storage period for the video data displayed on the viewing screen SC12 and sends it to the server S1. Based on the control command sent from the user terminals P1,..., the server S1 extends the storage period of the video data.
[0098] Viewing screen SC12A has the same configuration as viewing screen SC12. Therefore, in viewing screen SC12A shown in Figure 6, the same reference numerals are used for components that are the same as those in viewing screen SC12, and their explanation is omitted.
[0099] The viewing screen SC12A includes the captured image MV12, tag pin PN122, dropdown SL12, playback date and time information DT12, play button BT121, add comment button BT122, thumbnail images SM121, SM122, SM123, SM124, comment information CM121~CM124, digest video creation button BT123, and digest video creation button BT125.
[0100] Each of the thumbnail images SM121 to SM124 is an image captured at the time when a comment tag corresponding to each comment information CM121 to CM124 was assigned. When the processor 11 receives a tagging request (addition of a comment tag), it generates a comment tag. The processor 11 associates the generated comment tag with the comment time and the image captured at the time the user requested tagging, and sends it to the server S1. In this case, the video management system 100 may also store the thumbnail images SM121 to SM124 in the tag information database DB1, or it may extract the image captured at the time corresponding to each comment tag from the video data when generating the viewing screen SC12A, and generate a viewing screen SC12A that displays the extracted image captured images as thumbnail images SM121 to SM124.
[0101] The commented digest video creation button BT125 accepts an operation for generating a commented digest video in which predetermined sections before and after the positions (times) at which each comment tag is assigned are clipped, and comments are added as subtitles to a digest video obtained by connecting the clipped predetermined sections. When the commented digest video creation button BT125 is selected (pressed) by a user, the processor 11 generates a control command requesting generation of a commented digest video of the video data, and transmits the control command to the server S1.
[0102] The editing processing unit 212 of the server S1 clips a predetermined section including the time at which the comment tag is assigned (for example, several seconds before and after the time) based on a control command transmitted from the user terminal P1, ... . The editing processing unit adds subtitles based on the comment to the clipped video data, and generates a commented digest video by connecting the respective pieces of video data after subtitle addition along the time series of the comment tags.
[0103] As described above, the video management system 100 according to the embodiment performs imaging by the user terminals P1, ... that are mobile terminals, and thus can accept input of comments by a user after imaging even when it is difficult for the user to input comments in real time. The video management system 100 can support comment editing (input) by a user by displaying the tag pin PN122 at the comment assignment time with respect to a display section of video data, and the comment information CM121 to CM124 assigned to the display section.
[0104] Further, the video management system 100 according to the embodiment displays the comment information CM121 to CM124 assigned to a display section and the thumbnail images SM121 to SM124 respectively corresponding to the comment information CM121 to CM124, thereby making it possible to intuitively grasp what the content of a comment to be input is, and can support comment input by a user.
[0105] <Comment Generation Method Using AI (Trained Model)> Next, with reference to Figure 7, we will explain how to generate comments using AI (a pre-trained model). Figure 7 shows an example of the imaging screen SC11A and the AI viewing screen SC14. Note that the imaging screen SC11A and AI viewing screen SC14 shown in Figure 7 are just examples and are not limited to them.
[0106] When the application for performing imaging by the imaging unit 13 is launched on the user terminal P1, it generates an imaging screen SC11A and displays it on the display unit 14. Note that the imaging screen SC11A shown in Figure 7 has the same configuration as the imaging screen SC11. Therefore, in the imaging screen SC11A shown in Figure 7, the same reference numerals are used for components that are the same as those in the imaging screen SC11, and their explanation is omitted.
[0107] The AI button BT114 accepts operations to enable the detection process of scenes that are the target of comment generation from the captured image captured by the imaging unit 13, and the comment generation process based on the detection results, using deep learning with a pre-trained model.
[0108] When the AI button BT114 is selected (pressed) by the user, the processor 11 accepts user input to specify a trained model capable of detecting the scene to be detected. For example, the user selects one or more trained models from the following: "Chemical Leak" trained model for detecting situations such as hazardous substances overflowing from containers, machinery, etc., or a wet floor; "Abnormal Noise" trained model for detecting abnormal noises from machinery when it is overloaded; "Inappropriate Working Posture" trained model for detecting workers working in postures that put a strain on their waists, etc.; "Machine Malfunction" trained model for detecting situations such as smoke or sparks coming from machinery; and "Worker Collision or Walking with Collision Risk" trained model for detecting situations such as workers passing each other at close range where they are likely to collide. The processor 11 generates a control command requesting the transmission of the trained model selected by the user and sends it to the server S1.
[0109] The scenes detected by the trained model are not limited to the examples described above. For example, the processor 11 may use the trained model to detect scenes from the video data such as scenes with a lot of movement of objects in the captured image, scenes with a changed background, or detection scenes in which a person or object set as the detection target is detected.
[0110] Server S1 extracts the corresponding trained model from the AI database DB3 based on the control commands sent from user terminals P1,... and sends it to user terminals P1,....
[0111] User terminals P1,... acquire the trained model sent from server S1. When the user selects (presses) the imaging button BT113, the user terminals P1,... use the trained model to detect scenes (images) that are the target of comment generation from the images captured by the imaging unit 13. The processor 11 uses the trained model to generate the reason for the detection of the detected scene, or a comment that explains the detected scene. The processor 11 associates the comment tag, the comment time indicating the location (time) of the detected scene, the comment, the camera ID, the AI flag indicating that this comment was generated by the AI (trained model), and the AI type of the trained model used, and sends them to server S1.
[0112] Server S1 associates the comment tag (tag ID), comment time, comment, camera ID, AI flag, and AI type sent from processor 11 and stores them in the tag information database DB1.
[0113] After the AI button BT114 on the imaging screen SC11A is pressed and the imaging unit 13 has completed the imaging operation, the user terminal P1 receives an operation to specify the video data to be edited based on the user's operation, and obtains the specified video data and the tag information already attached to this video data from the server S1. Based on the obtained video data and tag information, the user terminal P1 generates an AI viewing screen SC14 that can accept the user's comment editing operation and displays it on the display unit 14.
[0114] Note that the AI viewing screen SC14 has the same configuration as viewing screens SC12 and SC12A. Therefore, in the AI viewing screen SC14 shown in Figure 7, the same components as viewing screens SC12 and SC12A are given the same symbols, and their explanation is omitted. Also, the AI viewing screen SC14 shown in Figure 7 displays only comments generated by a trained model as an example, but is not limited to this. The AI viewing screen SC14 may contain a mixture of comments generated by a trained model and comments entered by the user.
[0115] The AI viewing screen SC14 includes the captured image MV12, tag pin PN122, dropdown SL12, playback date and time information DT12, play button BT121, add comment button BT122, thumbnail images SM141, SM142, SM143, SM144, comment information CM141, CM142, CM143, CM144, and AI button BT115.
[0116] Each of the comment information entries CM141 to CM144 includes the comment time and a comment generated by the trained model, which explains the reason for the scene detection by the trained model or the detected scene. In the example shown in Figure 5, the comment information entries CM141 to CM144 include information about the location (time) where the comment tag was attached and the reason for the scene detection as the comment content. Specifically, comment information CM141 is a comment attached to the video data at the time "10:12:45", indicating that the reason for the scene detection is "Collision risk present". When the processor 11 receives a user operation on the area where the comment content is displayed, it accepts an edit operation for the comment content. When the processor 11 accepts a comment edit operation by the user, it deletes the corresponding comment information and the AI icon attached to the thumbnail image.
[0117] Each of the thumbnail images SM141 to SM144 is an image captured of a scene detected by the trained model, and is an image captured at the time when the comment tags corresponding to each of the comment information CM141 to CM144 were added. AI icons IC141, IC142, IC143, and IC144 are displayed nearby each of the thumbnail images SM141 to SM144 to indicate that it is a scene detected by the trained model.
[0118] Note that the AI button may consist of only the AI button BT115. In that case, first, when the user selects (presses) the imaging button BT113 on the user terminal P1, the imaging unit 13 starts imaging. Next, when the user selects (presses) the imaging button BT113 on the user terminal P1, the imaging unit 13 stops imaging. After the imaging is complete, the user terminal P1 generates the AI viewing screen SC14 and displays it on the display unit 14. At this time, the comment information is not yet displayed on the AI viewing screen SC14. When the user presses the AI button BT115 on the user terminal P1, the trained model generates comment information CM141 to CM144 and displays the generated comment information CM141 to CM144 on the AI viewing screen SC14.
[0119] As described above, the video management system 100 in this embodiment, by capturing images using user terminals P1, ... which are mobile terminals, can automatically generate detection of scenes detectable by the trained model and the reason for detecting these scenes, even when it is difficult for the user to input comments in real time, by accepting the user's operation to select a trained model. By performing scene detection using the trained model and generating the reason for detecting these scenes, the video management system 100 can reduce or eliminate the effort required for tagging and inputting or editing comments by the user.
[0120] <How to search for video data> Next, the method for searching for video data will be explained with reference to Figures 8 and 9, respectively. Figure 8 shows an example of the search results screen SC15. Figure 9 shows an example of the search results screen SC15A and the viewing screen SC15B. Note that the search results screens SC15, SC15A, and viewing screen SC15B shown in Figures 8 and 9 are examples and are not limited to them.
[0121] When a user terminal P1, ... receives a search operation for video data from the user, it generates a search results screen SC15 showing the search results for video data and displays it on the display unit 14. The search results screen SC15 includes search period input fields INP151 and INP152, a search button BT151, a search word input field INP153, and the search results video data MV151, MV152, MV153, and MV154, respectively.
[0122] The search period input fields INP151 and INP152 accept input operations for the period during which the video data to be searched was captured. The processor 11 accepts input operations from the user for the search start date and time in the search period input field INP151 and the search end date and time in the search period input field INP152.
[0123] The search button BT151 accepts a user's operation to search for video data. The processor 11 generates a control command requesting a search for video data captured during the period entered in the search period input fields INP151 and INP152, and sends it to the server S1. The server S1 extracts video data from the video storage DB2 that was captured during the period entered in the search period input fields INP151 and INP152, and sends it to the user terminals P1, ...
[0124] The search results screen SC15 shown in Figure 8 includes video data MV151 to MV154, respectively, captured between the period entered in the search period input fields INP151 and INP152, "2024 / 12 / 23 15:29" and "2025 / 01 / 16 14:26".
[0125] When a user terminal P1, ... receives a search operation for video data using keywords from the user, it generates a search results screen SC15A showing the search results for video data and displays it on the display unit 14. The search results screen SC15A includes search period input fields INP151 and INP152, a search button BT151, a search word input field INP153, and the search result video data MV151.
[0126] The search word input field INP153 accepts the user's input of keywords for searching video data. The processor 11 generates a control command requesting a search for video data captured during the period entered in the search period input fields INP151 and INP152, and which has comments attached that include the keyword entered in the search word input field INP153, and sends it to the server S1.
[0127] Server S1 extracts video data from the video storage that was captured during the period entered in the search period input fields INP151 and INP152, and that has comments containing the keyword entered in the search word input field INP153. Server S1 then sends the extracted video data to user terminals P1, ...
[0128] User terminals P1,... generate a search results screen SC15A containing video data MV151, which is the search result transmitted from server S1, and display it on the display unit 14. When user terminals P1,... receive a selection operation from the user for any of the video data MV151 displayed on the search results screen SC15A, they generate a viewing screen SC15B that makes the selected video data MV151 viewable and display it on the display unit 14. The viewing screen SC15B includes a display area AR15B, a search word input field INP157, and comment information CM151.
[0129] Display area AR15B is the area that displays (plays) the video data selected by the user. In Figure 9, display area AR15B is displaying (playing) the video data MV151 selected by the user.
[0130] The search word input field INP157 accepts a search operation for comments attached to video data. When one or more words or sentences are entered as keywords in the search word input field INP157, the processor 11 selects and displays comment information CM151 from the comments associated with the video data that contain the entered keywords. The viewing screen SC15B shown in Figure 9 detects the comment "This is a video for playback testing!" which contains the keyword "video" entered in the search word input field INP157, and displays this comment information CM151. Note that the viewing screen SC15B may be the same as the viewing screen SC12 and the AI viewing screen SC14. It may also be possible to edit comments from the viewing screen SC15B.
[0131] Furthermore, keywords entered in the search term input field INP157 may be given priority when searching or displaying search results.
[0132] As a result, the video management system 100 according to this embodiment can search for and view video data captured by each user terminal P1,.... Furthermore, the video management system 100 can improve the ease of searching for video data requested by users by using comments added by each user terminal P1,....
[0133] <How to generate comments using voice input> Next, the method for generating comments through user operation will be explained with reference to Figures 10 and 11, respectively. Figure 10 shows an example of the imaging screen SC16 and SC16A during voice input. Figure 11 shows an example of the viewing screen SC17 during voice input. Note that the imaging screen SC16, SC16A and viewing screen SC17 shown in Figures 10 and 11 are examples and are not limited thereto.
[0134] When the application for performing imaging by the imaging unit 13 is launched on the user terminal P1, it generates imaging screens SC16 and SC16A and displays them on the display unit 14. Note that the imaging screens SC16 and SC16A shown in Figure 10 have the same configuration as imaging screens SC11 and SC11A. Therefore, in the imaging screens SC16 and SC16A shown in Figure 10, the same reference numerals are used to indicate the same configuration as imaging screens SC11 and SC11A, and their explanation is omitted.
[0135] When the application for performing imaging by the imaging unit 13 is launched on the user terminal P1, it generates an imaging screen SC16 and displays it on the display unit 14. The imaging screen SC16 includes the captured image MV11, a mute button BT112, an imaging button BT113, and an audio comment button BT116.
[0136] The voice comment button BT116 accepts the operation of the control unit 15 to record the user's spoken voice. When the voice comment button BT116 is selected (pressed) by the user, the processor 11 starts recording the user's spoken voice from the control unit 15. When the voice comment button BT116 is selected by the user, it changes to a state (voice comment button BT117) that indicates that the user's spoken voice is being recorded.
[0137] Unlike the sound pickup unit 16, which picks up ambient sounds around the user terminals P1,..., the operation unit 15 primarily picks up the user's spoken voice to accept voice input of comments from the user. The processor 11 performs speech recognition on the voice picked up by the operation unit 15 and performs natural language processing on the speech recognition results to accept voice input of comments from the user.
[0138] In the example shown in Figure 10, the user selects (presses) the voice comment button BT116, then speaks VO161 "Please see the operation from here" and VO162 "I will finish the operation". The processor 11 performs speech recognition on VO161 and VO162, which are picked up by the operation unit 15, and generates the comments "Please see the operation from here" and "I will finish the operation" by performing natural language processing on the speech recognition results. The processor 11 generates a comment tag, using the timing when the voice comment button BT116 is pressed or when the utterance corresponding to the comment begins as the comment time. The processor 11 associates the generated comment tag, the comment time, and the comment and sends them to the server S1.
[0139] After the user terminal P1 receives the operation to end imaging from the imaging unit 13, if it receives an operation to specify video data to be edited based on the user's operation, it obtains the specified video data and the tag information already attached to this video data from the server S1. Based on the obtained video data and tag information, the user terminal P1 generates a viewing screen SC17 that can accept the user's comment editing operation and displays it on the display unit 14.
[0140] Note that the viewing screen SC17 shown in Figure 11 has the same configuration as viewing screens SC12 and SC12A. Therefore, in the viewing screen SC17 shown in Figure 11, the same reference numerals are used to indicate the same configuration as viewing screens SC12 and SC12A, and their explanation is omitted.
[0141] The viewing screen SC17 includes the captured image MV12, the tag pin PN122, the dropdown SL12, the playback date and time information DT12, the play button BT121, the comment add button BT122, and the comment information CM171, CM172, CM173, and CM174.
[0142] Each of the comment information entries CM171 to CM174 contains information about comments tagged with the current video data display section. In the example shown in Figure 11, comment information entries CM171 to CM174 include information about the location (time) where the comment tag was assigned and the comment content. Specifically, comment information CM172 indicates a comment assigned to the video data at the time "10:12:49" via voice input, with the comment content being "Please watch the operation from here." Comment information CM174 indicates a comment generated by operating the comment tag button BT111 or the comment add button BT122, with the comment content being empty.
[0143] As described above, the video management system 100 according to the embodiment can accept comments via the user's spoken voice, rather than manual comment input, even when it is difficult for the user to input comments in real time, by capturing images using user terminals P1,... which are mobile terminals. By accepting comments via the user's spoken voice, the video management system 100 can realize real-time comment input, and the effort required for tagging and comment input or editing by the user can be reduced or eliminated.
[0144] (Modified example of the embodiment) In the video management system 100 according to the embodiment, server S1 manages video data and comment information using a tag information database DB1 and a video storage DB2. In the following embodiment, server S1A of the video management system 100A will be described as managing video data and comment information using a tag information database DB1, a video storage DB2A, and a video management database DB4.
[0145] The video management system 100A, which is a modified version of the embodiment, has the same configuration and functions as the video management system 100 according to the embodiment. Therefore, in the following description of the embodiment, the same configuration and functions as the video management system 100 according to the embodiment will not be described.
[0146] The video management system 100 according to the embodiment will be described with reference to Figures 12 and 13, respectively. Figure 12 is a block diagram showing an example of the overall configuration of the video management system 100A according to a modified example of the embodiment. Figure 13 is a diagram showing an example of data storage for various types of data in a modified example of the embodiment.
[0147] The video management system 100A includes user terminals P1,... that accept imaging operations and comment addition operations from users, and a server S1A that manages the video images captured by P1,... and comments describing the content of the video, linking them to the user terminals. Note that the video management system 100A shown in Figure 12 is just one example and is not limited thereto.
[0148] Server S1A is connected to each of the multiple user terminals P1, ... via a network NW, enabling data communication. Server S1A manages video data captured by each of the multiple user terminals P1, ... and comments associated with the video data. Server S1A includes a communication unit 20, a processor 21A, a memory 22, a tag information database DB1, a video storage DB2A, and a video management database DB4.
[0149] The processor 21A is configured, for example, using a CPU or FPGA, and works in cooperation with the memory 22 to perform various processing and control operations. Specifically, the processor 21A refers to the programs and data held in the memory 22 and executes those programs to realize various functions such as the search unit 211 and the editing and processing unit 212. The functions described above are only a part of the functions that the processor 21A can realize, and are not limited to these.
[0150] The video storage DB2A is a so-called storage device, and is configured using a storage medium such as flash memory, HDD, or SSD. The video storage DB2A stores video data captured by multiple user terminals P1,... The video storage DB2A stores, for example, the TS file names of the files in which the video data is stored. The video storage DB2A shown in Figure 13 stores TS file names such as "xxxxxxx.Ts" and "yyyyyyy.Ts".
[0151] The video management database DB4 is a so-called storage system, configured using storage media such as flash memory, HDD, or SSD. The video management database DB4 stores information about video data captured by multiple user terminals P1,...
[0152] The video management database DB4 shown in Figure 13 stores information about video data indicating that the video data stored in the TS file name "xxxxxxxxx.Ts" has an acquisition start time of "2024 / 1 / 4 10:54:32.123", a video duration of "1.003", and the camera ID used for acquisition is "0001", and also information about video data stored in the TS file name "yyyyyyy.Ts" has an acquisition start time of "2024 / 1 / 4 10:57:45.456", a video duration of "1.003", and the camera ID used for acquisition is "0001".
[0153] When the processor 21A of server S1A extracts tag information (for example, time information of comment tags, or comment content, etc.) associated with each video data from the tag information database DB1 based on control commands sent from user terminals P1,..., it searches for and extracts tag information associated with the camera ID of user terminals P1,.... Here, each user terminal P1,... captures only one video data that a single camera can capture during the same time period. Based on the camera ID and the capture start time, the processor 21A searches and retrieves the video data that user terminals P1,... request to be transmitted, the TS file name in which the video data is stored, and the tag information (comments) attached to the video data from the tag information database DB1, video storage DB2A, and video management database DB4, respectively.
[0154] <Example of imaging procedure for a video management system> Next, with reference to Figure 14, an example of the imaging procedure of the video management system 100A in a modified example of the embodiment will be described. Figure 14 is a sequence diagram showing an example of the imaging procedure of the video management system 100A according to a modified example of the embodiment. Note that the example of the imaging procedure of the video management system 100A in a modified example of the embodiment includes the same procedure as the example of the imaging procedure of the video management system 100 shown in Figure 3. Therefore, in the following description of the example of the imaging procedure of the video management system 100A, the same reference numerals will be used for processes similar to those of the video management system 100 shown in Figure 3, and their explanations will be omitted.
[0155] Server S1A receives the image capture start tag sent from user terminal P1. Server S1A associates the image capture start tag with the image capture start time information and the camera ID and stores it in the tag information database DB1 (St13A).
[0156] In the imaging process LP11A, the server S1A generates a file name as identification information to enable the identification of video data transmitted from the user terminal P1 (St16). The server S1A stores the video data transmitted from the user terminal P1 in the video storage DB2, associating it with the file name (St17). The server S1A also stores the video data transmitted from the user terminal P1, associating the image capture start time, the capture time (video duration), the camera ID, and the file name generated by the server S1A, in the video management database DB4 (St17A). Note that the image capture start time in step St13A is the video recording start time, and the recording start time in steps St15 and St17A is the file recording start time.
[0157] As described above, the video management system 100A in the modified embodiment can accept a tagging operation for the location where a comment is added to the video data, that is, the imaging time at which the comment is added. This allows the video management system 100A to add a tag that makes it possible to refer to the imaging time at which a comment is requested after imaging, even if the user who is taking the image cannot input a comment in real time during imaging.
[0158] The video management systems 100 and 100A described above have been explained separately, but the technologies disclosed in this disclosure can be combined in any way. For example, the tagging process may be performed using a combination of selecting (pressing) the comment tag button BT111 and automatically assigning tags using a trained model (AI), and these tagging methods may also be partially combined with comment input using voice input.
[0159] (Note) The following technologies are disclosed based on the above description of embodiments.
[0160] (Technology 1) A video management method performed by a computer (user terminal P1,...) equipped with a camera (imaging unit 13), The camera (imaging unit 13) is made to perform imaging. During imaging by the camera (imaging unit 13), the system accepts user input at the timing when a comment is to be added, and generates a tag indicating the timing of the comment addition based on the user input. After the camera (imaging unit 13) has finished capturing images, the captured video data and the generated tag are displayed. The system accepts the input operation for the aforementioned comment for each tag, links the comment with the video data captured by the camera (imaging unit 13), and registers it in a storage medium (tag information database DB1, video storage DB2, or video management database DB4, etc.). Video management methods. As a result, user terminals P1,... in this disclosure can accept tagging operations for the location where comments are added to the video data, that is, the capture time when comments are added. As a result, even if the user who is taking the image cannot input comments in real time during capture, user terminals P1,... can add tags that allow them to refer to the capture time when comment input is requested after the image is taken, and can accept comment input corresponding to the time each tag was added after the image is taken.
[0161] (Technology 2) The system displays the tag, the captured image taken at the time the tag was generated, and the time the tag was generated. The video management method described in (Technical 1). As a result, user terminals P1,... in this disclosure can facilitate the recall of comments to be added at the time the captured image was taken, and assist in the input of comments, by displaying the captured image.
[0162] (Technology 3) The imaging screens SC11 and SC11A, which include the image captured by the camera (imaging unit 13), are displayed. The imaging screen includes a tag button (comment tag button BT111) that can accept the tag assignment operation. The video management method described in (Technical 1). As a result, user terminals P1,... in this disclosure can display captured images captured in real time by the imaging unit 13 during imaging, and accept tagging operations on the displayed captured images. Therefore, the user can perform imaging and tagging operations simultaneously by checking the displayed imaging screen and selecting the comment tag button BT111 displayed on the same screen when they wish to add a comment.
[0163] (Technology 4) A predetermined section of the video data, including the time before and after the time when the tag was generated, is extracted for each tag. It generates a digest video by connecting the extracted video clips in chronological order. The video management method described in (Technical 1). As a result, user terminals P1,... in this disclosure can generate digest videos that allow users to efficiently view only the sections of video data that they wish to watch, when sharing video data with multiple different users.
[0164] (Technology 5) Based on the comments corresponding to the aforementioned extracted video footage, subtitles are generated. A digest video is generated by adding the subtitles to the aforementioned extracted video footage. The video management method described in (Technical 4). As a result, user terminals P1,... in this disclosure can efficiently view only the sections of video data that should be viewed when sharing video data with multiple different users, and can also generate a digest video with comment content displayed as subtitles.
[0165] (Technology 6) Using a trained model capable of detecting scenes to which the tag is assigned from images captured by the camera (imaging unit 13), the detection of the scene and the generation of the tag at the time the scene is detected are performed. The video management method described in (Technical 1). As a result, user terminals P1,... in this disclosure can eliminate the effort required for users to manually add tags, and, when used in conjunction with the comment tag button BT111, can suppress users from forgetting to add tags.
[0166] (Technology 7) The user operation described above accepts the input of search words (keywords), The system retrieves video data from the storage medium that includes comments containing the entered search term and displays it. The video management method described in (Technical 1). As a result, user terminals P1,... in this disclosure can acquire video data that the user requests to view based on comments and display it in a viewable format.
[0167] (Technology 8) Based on the user operation, the storage period for storing the video data in the storage medium is extended. The video management method described in (Technical 1). As a result, user terminals P1,... in this disclosure can extend the storage period of each video data based on the user's needs.
[0168] (Technology 9) Multiple video management devices (user terminals P1, ...) each having a camera (imaging unit 13), A system (video management system 100, 100A) comprising servers S1, S1A for managing video data captured by the plurality of video management devices, the video management device (user terminal P1, ...) The aforementioned camera (imaging unit 13), An operating unit 15 capable of receiving user input, The system includes a display unit 14 that displays video data captured by the camera (imaging unit 13), The operation unit 15 accepts user input at the timing when the user wishes to add a comment while the camera (imaging unit 13) is capturing images, and generates a tag indicating the timing of the comment addition based on the user input. The display unit 14 displays the captured video data and the generated tag after the camera (imaging unit 13) has finished capturing images. The operation unit 15 receives the input operation for each tag, associates the comment with the video data captured by the camera (imaging unit 13), and registers it with the servers S1 and S1A. Video management device (user terminal P1, ...). As a result, user terminals P1,... in this disclosure can accept tagging operations for the location where comments are added to the video data, that is, the capture time when comments are added. As a result, even if the user who is taking the image cannot input comments in real time during capture, user terminals P1,... can add tags that allow them to refer to the capture time when comment input is requested after the image is taken, and can accept comment input corresponding to the time each tag was added after the image is taken.
[0169] (Technology 10) Multiple terminals (user terminals P1, ...) each having a camera (imaging unit 13), A video management system 100, 100A comprising servers S1, S1A for managing video data captured by the aforementioned multiple terminals (user terminals P1, ...), The aforementioned terminals (user terminals P1, ...) The camera (imaging unit 13) is instructed to perform imaging, and the captured video data is transmitted to the servers S1 and S1A. During imaging by the camera (imaging unit 13), the system accepts user input at the timing when a comment is to be added, generates a tag indicating the timing of the comment addition based on the user input, and sends it to the servers S1 and S1A. The aforementioned servers S1 and S1A are The video data transmitted from the aforementioned terminal (user terminal P1, ...) and the aforementioned tag are linked and stored. After the imaging by the terminal (user terminal P1,...) is completed, the captured video data and the tag corresponding to the video data are transmitted to the terminal (user terminal P1,...). The aforementioned terminals (user terminals P1, ...) The aforementioned video data and the tags corresponding to the aforementioned video data are displayed. The input operation for the aforementioned comment is accepted for each of the aforementioned tags, the comment is linked to the video data and sent to the servers S1 and S1A. The aforementioned servers S1 and S1A are The aforementioned comments and the aforementioned video data are linked and stored. Video management system 100, 100A. As a result, the video management systems 100 and 100A in this disclosure can accept tagging operations for the location where comments are added to the video data, that is, the capture time when comments are added. As a result, even if the user who is taking the image cannot input comments in real time during capture, the video management systems 100 and 100A can add tags that allow them to refer to the capture time when comment input is requested after capture, and can accept comment input corresponding to the time when each tag was added after capture.
[0170] (Technology 11) A video management program performed by a processor 11 that is communicatively connected to a camera (imaging unit 13), The steps include causing the camera (imaging unit 13) to perform imaging, The camera (imaging unit 13) takes an image while the user wants to add a comment, and based on the user's input, it generates a tag indicating the timing of the comment addition. After the camera (imaging unit 13) has finished capturing images, the steps include displaying the captured video data and the generated tags, To implement the steps of receiving the input operation for the aforementioned comment for each tag, linking the comment with the video data captured by the camera (imaging unit 13), and registering it in a storage medium (tag information database DB1, video storage DB2, or video management database DB4, etc.), Video management program. As a result, user terminals P1,... in this disclosure can accept tagging operations for the location where comments are added to the video data, that is, the capture time when comments are added. As a result, even if the user who is taking the image cannot input comments in real time during capture, user terminals P1,... can add tags that allow them to refer to the capture time when comment input is requested after the image is taken, and can accept comment input corresponding to the time each tag was added after the image is taken.
[0171] Although various embodiments have been described above with reference to the drawings, it goes without saying that this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the various embodiments described above can be combined arbitrarily without departing from the spirit of the invention. [Industrial applicability]
[0172] This disclosure is useful as a video management method, video management device, video management system, and video management program that can more easily generate video data with comments attached. [Explanation of Symbols]
[0173] 10,20 Communications Department 11,21,21A processor 12.22 memory 13 Imaging Unit 14 Display section 15 Control section 16 Sound collection section 100,100A Video Management System 111,211 Search section 212 Editing and Processing Department BT111 Comment Tag Button BT112 Mute Button BT113 Image capture button BT114, BT115 AI button BT116, BT117 Voice Comment Button BT121 Play button BT122 Add Comment Button BT123 Digest Video Creation Button BT124 Storage Period Extension Button BT125 Commentary Digest Video Creation Button BT151 Search Button CM121, CM122, CM123, CM124, CM141, CM142, CM143, CM144, CM151, CM171, CM172, CM173, CM174 Comment Information DB1 Tag Information Database DB2, DB2A Video Storage DB3 AI Database DB4 Video Management Database INP153, INP157 Search word input field MV11, MV12 captured images NW Network P1 User Terminal S1, S1A Server SC11, SC11A, SC16, SC16A imaging screen SC12, SC12A, SC15B, SC17 Viewing Screen SC14 AI viewing screen SC15, SC15A Search Results Screen SM121, SM122, SM123, SM124, SM141, SM142, SM143, SM144 Thumbnail Images VO161, VO162 utterances
Claims
1. A video management method performed by a computer equipped with a camera, The camera is made to perform imaging, During image capture by the aforementioned camera, the system accepts user input at the time when the user wishes to add a comment, and generates a tag indicating the timing of the comment addition based on the user input. After the camera has finished capturing images, the captured video data and the generated tag are displayed. The system accepts the input operation for the aforementioned comment for each tag, and links the comment with the video data captured by the camera and registers it in the storage medium. Video management methods.
2. The system displays the tag, the captured image taken at the time the tag was generated, and the time the tag was generated. The video management method according to claim 1.
3. The camera displays an image capture screen including the captured image captured by the aforementioned camera. The imaging screen includes a tag button capable of accepting the tag assignment operation. The video management method according to claim 1.
4. A predetermined section of the video data, including the time before and after the time when the tag was generated, is extracted for each tag. It generates a digest video by connecting the extracted video clips in chronological order. The video management method according to claim 1.
5. Based on the comments corresponding to the aforementioned extracted video footage, subtitles are generated. A digest video is generated by adding the subtitles to the aforementioned extracted video footage. The video management method according to claim 4.
6. Using a trained model capable of detecting scenes to which the tag is assigned from images captured by the aforementioned camera, the detection of the scene and the generation of the tag at the time the scene is detected are performed. The video management method according to claim 1.
7. The user operation described above accepts the input of a search term. The system retrieves video data from the storage medium that includes comments containing the entered search term and displays it. The video management method according to claim 1.
8. Based on the user operation, the storage period for storing the video data in the storage medium is extended. The video management method according to claim 1.
9. Multiple video management devices equipped with cameras, The video management device of a system comprising a server that manages video data captured by the plurality of video management devices, The aforementioned camera, An operating unit that can accept user input, The system includes a display unit that displays video data captured by the aforementioned camera, The operation unit accepts user input at the timing when the user wishes to add a comment while the camera is capturing images, and generates a tag indicating the timing of the comment addition based on the user input. The display unit displays the captured video data and the generated tag after the camera has finished capturing images. The operation unit receives the input operation for the comment for each tag, associates the comment with the video data captured by the camera, and registers it on the server. Video management device.
10. Multiple devices equipped with cameras, A video management system comprising a server that manages video data captured by the aforementioned multiple terminals, The aforementioned terminal is The camera is made to perform imaging, and the captured video data is transmitted to the server. During image capture by the aforementioned camera, the system accepts user input at the time when the user wishes to add a comment, generates a tag indicating the timing of the comment addition based on the user input, and sends it to the server. The aforementioned server, The video data transmitted from the terminal and the tag are linked and stored together. After the terminal has finished capturing images, the captured video data and the tag corresponding to the video data are transmitted to the terminal. The aforementioned terminal is The aforementioned video data and the tags corresponding to the aforementioned video data are displayed. The input operation for the aforementioned comment is accepted for each of the aforementioned tags, the comment is linked to the video data and sent to the server, The aforementioned server, The aforementioned comments and the aforementioned video data are linked and stored. Video management system.
11. A video management program performed by a processor that is communicatively connected to a camera, The steps include causing the aforementioned camera to perform imaging, The steps include: accepting user input at the time when the user wishes to add a comment while the camera is capturing images; and generating a tag indicating the timing of the comment addition based on the user input; After the camera has finished capturing images, the steps include displaying the captured video data and the generated tags, To implement the steps of receiving the input operation for the aforementioned comment for each of the aforementioned tags, and registering the comment and the video data captured by the camera in a storage medium, Video management program.
Citation Information
Patent Citations
Terminal device, server device, and program for recording work state by means of image
WO2016143749A1