Program, method, information processing device, and system
The system improves video editing convenience by analyzing captured images with multiple models, presenting detection markers, and generating edited videos, addressing the challenge of handling multiple tracking targets in surveillance systems.
Patent Information
- Application Number
- PCT/JP2025/001622
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2025-01-20
- Publication Date
- 2025-07-31
AI Technical Summary
Existing surveillance camera systems face challenges in improving convenience when editing moving images, particularly in handling multiple tracking targets within a frame.
A system and method that utilizes a computer program to analyze captured moving images using multiple pre-stored analysis models, present detection markers on a timeline, and generate edited videos based on user instructions, allowing for intuitive and efficient video editing.
Enhances the convenience of editing moving images by enabling users to intuitively grasp and efficiently process objects within the video, ensuring no missed signs before or after the object appears or exits.
Smart Images

Figure JP2025001622_31072025_PF_FP_ABST
Abstract
Description
Program, method, information processing device, and system
[0001] The present disclosure relates to a program, a method, an information processing device, and a system.
[0002] Patent Literature 1 describes a technology applicable to a surveillance camera system for designating a tracking target before or during tracking of an object. In Patent Literature 1, a user selects an object (tracking target) that the user wants to enlarge from among extracted tracking target candidates, and obtains a desired enlarged image (zoomed image).
[0003] JP 2009-251940 A
[0004] In Patent Document 1, one type of tracking target is extracted from one frame. However, one frame may contain multiple types of tracking targets. When editing videos based on images captured by a surveillance camera or the like, there is a demand for improved convenience.
[0005] An object of the present disclosure is to improve convenience when editing videos based on images captured by a surveillance camera or the like.
[0006] A program to be executed by a computer having a processor and a memory causes the processor to execute the following steps: acquiring a captured video, analyzing the acquired video using a plurality of pre-stored analytical models, each of which has been trained to detect an object, presenting to a user detection markers that indicate the timing at which the object is detected by the analytical models and that are associated with a timeline associated with the video, and generating an edited video from the captured video in response to a user instruction based on the detection markers.
[0007] According to the present disclosure, convenience can be improved when editing videos based on images captured by a surveillance camera or the like.
[0008] 1 is a block diagram showing the overall configuration of the system. FIG. 2 is a block diagram showing the configuration of the terminal device shown in FIG. 1. FIG. 3 is a diagram showing the functional configuration of the server. FIG. 4 is a diagram showing the data structure of an imaging device DB. FIG. 5 is a diagram showing the data structure of a detection log DB. FIG. 6 is a flowchart showing the operation when information detected by image analysis is recorded in a log. FIG. 7 is a schematic diagram showing an editing screen displayed on the display of the terminal device. FIG. 8 is a schematic diagram showing a part of the editing screen displayed on the display of the terminal device. FIG. 9 is a schematic diagram showing another example of the editing screen displayed on the display of the terminal device. FIG. 10 is a block diagram showing the basic hardware configuration of a computer 90.
[0009] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the following description, the same components are denoted by the same reference numerals. The names and functions of the components are also the same. Therefore, detailed descriptions thereof will not be repeated.
[0010] Overview: The system according to this embodiment analyzes captured video using multiple image analysis models. The system presents objects that allow users to understand which model detected the target object, and creates an edited video based on the object designation.
[0011] <1. Overall System Configuration Diagram> Fig. 1 is a block diagram showing an example of the overall configuration of a system 1. The system 1 shown in Fig. 1 includes, for example, a terminal device 10, a server 20, an imaging device 31, a sensor 32, and a speaker 33. The terminal device 10, the server 20, and the speaker 33 are communicatively connected via, for example, a network 80.
[0012] 1 shows an example in which the system 1 includes two terminal devices 10, but the number of terminal devices 10 included in the system 1 is not limited to two. The terminal devices 10 are terminals carried by camera users. The number of terminal devices 10 included in the system 1 may be less than three, or may be three or more.
[0013] In this embodiment, a collection of multiple devices may be considered as one server. The allocation of multiple functions required to realize the server 20 according to this embodiment to one or more pieces of hardware can be determined appropriately in consideration of the processing capacity of each piece of hardware and / or the specifications required for the server 20.
[0014] 1 may be, for example, a mobile terminal such as a smartphone or a tablet, a desktop personal computer (PC), a laptop PC, or a wearable terminal such as a head mounted display (HMD) or a wristwatch terminal.
[0015] The terminal device 10 includes a communication interface (IF) 12 , an input device 13 , an output device 14 , a memory 15 , a storage 16 , and a processor 19 .
[0016] The communication IF 12 is an interface for inputting and outputting signals so that the terminal device 10 can communicate with devices in the system 1, such as the server 20, for example.
[0017] The input device 13 is a device for receiving input operations from a user (for example, a touch panel, a touch pad, a pointing device such as a mouse, a keyboard, etc.).
[0018] The output device 14 is a device (such as a display or speaker) for presenting information to the user.
[0019] The memory 15 is for temporarily storing programs and data to be processed by the programs, and is a volatile memory such as a DRAM (Dynamic Random Access Memory).
[0020] The storage 16 is for storing data, and is, for example, a flash memory or a hard disk drive (HDD).
[0021] The processor 19 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.
[0022] The server 20 is realized by, for example, a computer connected to a network 80. As shown in FIG. 1, the server 20 includes a communication IF 22, an input / output IF 23, a memory 25, a storage 26, and a processor 29.
[0023] The communication IF 22 is an interface for inputting and outputting signals so that the server 20 can communicate with devices in the system 1, such as the terminal device 10, for example.
[0024] The input / output IF 23 functions as an interface with an input device for receiving input operations from the user and an output device for presenting information to the user.
[0025] The memory 25 is for temporarily storing programs and data to be processed by the programs, and is a volatile memory such as a DRAM.
[0026] The storage 26 is for storing data, and is, for example, a flash memory or a HDD.
[0027] The processor 29 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.
[0028] The image capturing device 31 is a device that receives light using a light receiving element and outputs the light as an image signal. The image capturing device 31 captures images at a frame rate that allows a series of movements to be recognized as a video. A frame rate that allows a series of movements to be recognized as a video is, for example, about 30 fps. The image capturing device 31 may be, for example, a camera that can capture a wide range of images covering 360 degrees. In this case, the image capturing device 31 is realized by, for example, a camera with an ultra-wide-angle lens or a fisheye lens.
[0029] The image capturing device 31 is installed at a position where it can see the entire site without any obstructions. If one image capturing device 31 cannot capture the entire site, multiple image capturing devices 31 are installed. When multiple image capturing devices 31 are installed, for example, a section of the site to be captured by each image capturing device 31 is set in advance. The image capturing device 31 outputs the acquired image signal to the server 20.
[0030] 2 is a block diagram showing an example configuration of the terminal device 10 shown in FIG. 1. The terminal device 10 shown in FIG. 2 is realized by a mobile terminal, a PC, or a wearable terminal. As shown in FIG. 2, the terminal device 10 includes a communication unit 120, an input device 13, an output device 14, an audio processing unit 17, a microphone 171, a speaker 172, a camera 161, a position information sensor 150, a storage unit 180, and a control unit 190. The blocks included in the terminal device 10 are electrically connected by, for example, a bus or the like.
[0031] The communication unit 120 performs processing such as modulation and demodulation for the terminal device 10 to communicate with other devices. The communication unit 120 performs transmission processing on signals generated by the control unit 190 and transmits the signals to an external device (for example, the server 20). The communication unit 120 performs reception processing on signals received from an external device and outputs the signals to the control unit 190.
[0032] The input device 13 is a device for inputting instructions or information by a user operating the terminal device 10. The input device 13 is realized, for example, by a touch-sensitive device 131 or the like, which inputs instructions by touching the operation surface. If the terminal device 10 is a PC or the like, the input device 13 may be realized by a reader, keyboard, mouse, or the like. The input device 13 converts instructions input by the user into electrical signals and outputs the electrical signals to the control unit 190. Note that the input device 13 may include, for example, a receiving port that receives electrical signals input from an external input device.
[0033] The output device 14 is a device for presenting information to a user operating the terminal device 10. The output device 14 is realized, for example, by a display 141. The display 141 displays data according to the control of the control unit 190. The display 141 is realized, for example, by an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.
[0034] The audio processing unit 17 performs, for example, digital-to-analog conversion processing of an audio signal. The audio processing unit 17 converts a signal provided from the microphone 171 into a digital signal and provides the converted signal to the control unit 190. The audio processing unit 17 also provides the audio signal to the speaker 172. The audio processing unit 17 is realized, for example, by a processor for audio processing. The microphone 171 receives audio input and provides an audio signal corresponding to the audio input to the audio processing unit 17. The speaker 172 converts the audio signal provided from the audio processing unit 17 into audio and outputs the audio to the outside of the terminal device 10.
[0035] The camera 161 is a device that receives light with a light receiving element and outputs the light as an image signal.
[0036] The position information sensor 150 is a sensor that detects the position of the terminal device 10, and is, for example, a GPS (Global Positioning System) module. The GPS module is a receiving device used in a satellite positioning system. The satellite positioning system receives signals from at least three or four satellites and detects the current position of the terminal device 10 equipped with the GPS module based on the received signals. The position information sensor 150 may detect the current position of the terminal device 10 from the position of the wireless base station to which the terminal device 10 is connected.
[0037] The storage unit 180 is realized by, for example, the memory 15, the storage 16, etc., and stores data and programs used by the terminal device 10. The storage unit 180 stores, for example, user information 181.
[0038] The user information 181 includes, for example, information about the user who uses the terminal device 10. The information about the user includes, for example, information for identifying the user, the user's name, age, address, date of birth, telephone number, email address, etc.
[0039] The control unit 190 is realized by the processor 19 reading a program stored in the storage unit 180 and executing instructions included in the program. The control unit 190 controls the operation of the terminal device 10. By operating in accordance with the program, the control unit 190 fulfills the functions of an operation reception unit 191, a transmission / reception unit 192, and a presentation control unit 193.
[0040] Operation reception unit 191 performs processing for receiving instructions or information input from input device 13. Specifically, for example, operation reception unit 191 receives information based on instructions input from touch-sensitive device 131 or the like. The instructions input via touch-sensitive device 131 or the like are, for example, editing instructions from the user.
[0041] The operation acceptance unit 191 also accepts voice instructions input from the microphone 171. Specifically, for example, the operation acceptance unit 191 receives a voice signal that is input from the microphone 171 and converted into a digital signal by the voice processing unit 17. The operation acceptance unit 191 acquires instructions from the user, for example, by analyzing the received voice signal and extracting predetermined nouns.
[0042] The transmitting / receiving unit 192 performs processing for the terminal device 10 to transmit and receive data to and from external devices such as the server 20 in accordance with a communication protocol. Specifically, for example, the transmitting / receiving unit 192 transmits editing instructions input by the user to the server 20. In addition, the transmitting / receiving unit 192 receives information about the user from the server 20.
[0043] The presentation control unit 193 controls the output device 14 to present information provided from the server 20 to the user. Specifically, for example, the presentation control unit 193 causes the information transmitted from the server 20 to be displayed on the display 141. In addition, the presentation control unit 193 causes the information transmitted from the server 20 to be output from the speaker 172.
[0044] 3 is a diagram showing an example of the functional configuration of the server 20. As shown in FIG. 3, the server 20 functions as a communication unit 201, a storage unit 202, and a control unit 203.
[0045] The communication unit 201 performs processing for the server 20 to communicate with external devices.
[0046] The storage unit 202 includes, for example, an imaging device database (DB) 2021 and a detection log database (DB) 2022 .
[0047] The image capturing device DB 2021 is a database for storing information about cameras installed for image capturing, as will be described in detail later.
[0048] The detection log DB 2022 is a database for storing information about objects detected by image analysis, etc. Details will be described later.
[0049] The first trained model 2023 is a model generated by having a machine learning model perform machine learning in accordance with a model learning program. The first trained model 2023 is, for example, a parameterized composite function formed by combining multiple functions that performs predetermined inference based on input data. The parameterized composite function is defined by a combination of multiple adjustable functions and parameters. The trained model according to this embodiment may be any parameterized composite function that meets the above requirements. The same applies to the second trained model 2024 and the third trained model 2025.
[0050] For example, when the first trained model 2023 is generated using a forward propagation type multi-layer network, the parameterized composite function is defined as a combination of, for example, a linear relationship between each layer using a weight matrix, a nonlinear relationship (or a linear relationship) using an activation function in each layer, and a bias. The weight matrix and bias are called parameters of the multi-layer network. The form of the parameterized composite function as a function changes depending on how the parameters are selected. In a multi-layer network, by appropriately setting the constituent parameters, it is possible to define a function that can output desirable results from the output layer. The same applies to the second trained model 2024 and the third trained model 2025.
[0051] The multi-layer network according to the present embodiment may be, for example, a deep neural network (DNN), which is a multi-layer neural network that is the subject of deep learning. As the DNN, for example, a convolution neural network (CNN) that targets images may be used.
[0052] The first trained model 2023 is a model that learns to output information about a human based on an input image when the image is input. The first trained model 2023 may be, for example, different trained models that are trained with different learning data depending on the information to be output.
[0053] For example, the first trained model 2023 is trained to output information about human attributes such as gender, age, race, and occupational status when an image is input. Specifically, in this embodiment, the first trained model 2023 identifies human attributes based on information about the physique, face, hairstyle, clothing, and belongings of the human included in the image. Furthermore, for example, the first trained model 2023 may be trained to output information about a specific person when an image is input. Specifically, in this embodiment, the first trained model 2023 may identify a specific person based on information about the face, hairstyle, and physique of the specific person included in the image. Furthermore, for example, the first trained model 2023 may be trained to output information about human body parts when an image is input. Specifically, in this embodiment, the first trained model 2023 may identify human body parts such as hands or a face based on the skeletal arrangement of the human included in the image. In this case, the first trained model 2023 identifies human body parts included in the image using a method such as open pose, for example.
[0054] In this case, the learning data may be, for example, a human image as input data, and the judgments about the attributes, specific persons, and human body parts included in the input data may be used as correct output data. In this case, the learning data may not include correct output data.
[0055] The second trained model 2024 is a model trained to output information about an object based on an input image when the image is input. The object here refers to an animal or plant, clothing, machinery, a vehicle, a craft, interior decoration, a building, etc. The second trained model 2024 may be, for example, a separate trained model trained with separate learning data depending on the information to be output.
[0056] For example, the second trained model 2024 learns to output information about an object when an image is input. Specifically, in this embodiment, the second trained model 2024 identifies the object based on information such as the size, shape, material, color, and accessories of the object included in the image.
[0057] In this case, the learning data may be, for example, an image of an object as input data, and a judgment about the object represented in this input data may be the correct output data. In this case, the learning data may not include the correct output data.
[0058] The third trained model 2025 is a model that learns to output information about a movement based on an input image when the image is input. The third trained model 2025 may be, for example, different trained models that are trained with different learning data depending on the information to be output.
[0059] For example, the third trained model 2025 learns to output information related to shopping when an image is input. Specifically, in this embodiment, the third trained model 2025 is trained to output that a shopping action has been performed when an image including the following content is input: - A customer taking out money or a card near a cash register in a retail store in an action spanning multiple frames - A customer picking up a product in a retail store in an action spanning multiple frames
[0060] In this case, the learning data may be, for example, an image of the inside of a retail store as input data, and the judgment regarding the behavior represented by this input data may be the correct output data. In this case, the learning data may not include the correct output data.
[0061] The control unit 203 is realized by the processor 29 reading a program stored in the storage unit 202 and executing instructions included in the program. By operating in accordance with the program, the control unit 203 performs functions shown as a reception control module 2031, a transmission control module 2032, an image analysis module 2034, an image editing module 2035, and a presentation module 2036.
[0062] The reception control module 2031 controls the process by which the server 20 receives signals from external devices in accordance with a communication protocol.
[0063] The transmission control module 2032 controls the process in which the server 20 transmits signals to external devices in accordance with a communication protocol.
[0064] The image analysis module 2034 analyzes images captured by the imaging device 31 and outputs information about objects at the site. The image analysis module 2034 inputs the images to the first learned model 2023, the second learned model 2024, and the third learned model 2025. More specifically, the image analysis module 2034 analyzes the image using, for example, the first learned model 2023, and after the analysis by the first learned model 2023 is completed, analyzes the image using the second learned model 2024. After the analysis by the second learned model 2024 is completed, the image analysis module 2034 analyzes the image using the third learned model 2025. The first learned model 2023, the second learned model 2024, and the third learned model 2025 output information about objects corresponding to each model in accordance with the input image. The image analysis module 2034 outputs information output from the first trained model 2023, the second trained model 2024, and the third trained model 2025 to the presentation module 2036. The image analysis module 2034 does not necessarily use the first trained model 2023, the second trained model 2024, or the third trained model 2025, and may analyze an image using a trained model specified by a user. The image analysis module 2034 may input an image with a limited analysis range to the trained model. The analysis range may be specified, for example, by the user. Furthermore, time may be added to the conditions for analyzing an image. For example, the image analysis module 2034 analyzes an image only if an object continues to be detected in the analysis range for a certain period of time or more.
[0065] The image editing module 2035 receives editing operations from the user and edits the video analyzed by the image analysis module 2034. Specifically, for example, the image editing module 2035 edits the video from the perspective of time, objects, range on the image, etc. The image editing module 2035 may edit the video in accordance with point-by-point operations by the user, or may automatically edit the video based on rules set by the user.
[0066] The presentation module 2036 presents information about objects in the field to the user via the terminal device 10. The presentation module 2036 may present information to the user based on an instruction from the user.
[0067] 4 and 5 are diagrams showing the data structure of the database stored in the server 20. Note that Fig. 4 and Fig. 5 are merely examples and do not exclude data not shown.
[0068] Fig. 4 is a diagram showing the data structure of the camera device DB 2021. As shown in Fig. 4, each record of the camera device DB 2021 includes, for example, an item "camera device ID," an item "coordinates," an item "address," an item "model," an item "installation date," an item "removal date," an item "operation status," an item "latest analysis," an item "current analysis status," an item "analysis frequency," and an item "captured video."
[0069] The item “camera device ID” indicates identification information for identifying the camera device 31 .
[0070] The item "coordinates" indicates the installation location of the camera device 31. Specifically, the item "coordinates" indicates the latitude and longitude of the installation location of the camera device 31. Note that the installation location of the camera device 31 may be determined by indicators other than coordinates.
[0071] The item "Address" indicates the address of the location where the image capturing device 31 is installed.
[0072] The item “model” indicates the model of the imaging device 31 .
[0073] The item "Installation Date" indicates the date on which the imaging device 31 was installed.
[0074] The item "Removal Date" indicates the date on which the image capturing device 31 was removed.
[0075] The item "operation status" indicates the operation status of the imaging device 31. Specifically, the item "operation status" is expressed as "in operation," "stopped," "out of order," or "removed" according to the status of the imaging device 31.
[0076] The item "Latest Analysis" indicates the time point of the most recent image that has been analyzed by the server 20. Specifically, the server 20 analyzes images received from the imaging device 31 at a predetermined timing, and the item "Latest Analysis" indicates the time point of the most recent image that has been analyzed. In other words, images up to the time point indicated by the item "Latest Analysis" in Figure 4 (e.g., 2023 / 10 / 22 23:59:59) have been analyzed.
[0077] The item "Current Analysis Status" indicates the time division of the image currently being analyzed by the server 20. The time division can be set arbitrarily. The divisions are, for example, as follows: Daily (00:00:00 to 23:59:59) Morning (00:00:00 to 9:59:59), Daytime (10:00:00 to 17:59:59), Nighttime (18:00:00 to 23:59:59)
[0078] The item "analysis frequency" indicates the frequency with which the server 20 analyzes images.
[0079] The item "video" indicates a video captured by the camera device 31. Specifically, the video is video data captured by the camera device 31. The video may also store reference information to a video data file located elsewhere.
[0080] 5 is a diagram showing the data structure of the detection log DB 2022. As shown in Fig. 5, each record in the detection log DB 2022 includes, for example, an item "camera ID," an item "detection model," an item "detection category," an item "detection frame," and an item "image area." When an object is detected by the trained model, a new record is created in the detection log DB 2022, and each piece of information is stored in the corresponding item.
[0081] The item “camera device ID” indicates identification information for identifying the camera device 31 .
[0082] The item "detection model" indicates the trained model that detected the object. Specifically, the first trained model 2023, the second trained model 2024, the third trained model 2025, etc. stored in the storage unit 202 are used as the detection model.
[0083] The item "Detection Category" indicates the category of the detected object. Specifically, for example, the category is a classification of the object detected by the trained model. For example, the object detected by the first trained model 2023 is classified as an "adult male" among humans.
[0084] The item "detection frame" indicates the frame of the image in which the object was detected.
[0085] The item "image area" indicates the area of the object on the image. Specifically, for example, when the image is divided into meshes, the item "image area" indicates the mesh in which the object appears on the image.
[0086] <3 Operation> The operation of the server 20 for analyzing images will be described. Although a retail store will be used as an example below, the application of this embodiment is not limited to retail stores. For example, this embodiment may be used to check rule enforcement in a hospital or for safety management in a factory.
[0087] The camera device 31 is attached to the ceiling, for example, from which the equipment in the store can be viewed. In this embodiment, the equipment is, for example, items and locations used for purchasing or selling. In this embodiment, the equipment includes, for example, shelves, entrances and exits, cash registers, and shopping carts. The camera device 31 captures images of the store at a frame rate that allows a series of movements to be recognized as a video. A frame rate that allows a series of movements to be recognized as a video is, for example, about 30 fps. The camera device 31 transmits image data to the server 20.
[0088] The frequency of transmission is a fixed cycle based on the data volume of the unsent video. For example, the image capture device 31 transmits the unsent video to the server 20 every time the data volume of the unsent video reaches a predetermined volume. The frequency of transmission may also be a fixed cycle based on the time of the unsent video. For example, the image capture device 31 transmits the unsent video to the server 20 every time a predetermined time has elapsed since the most recent transmission. Alternatively, the image capture device 31 may immediately transmit the captured video to the server 20.
[0089] FIG. 6 is a flowchart showing the operation when information detected by image analysis is recorded in a log.
[0090] In step S61, the control unit 203 analyzes an image captured by the imaging device 31 using the first trained model 2023 recorded in the storage unit 202. Specifically, the image analysis module 2034, for example, inputs image data transmitted from the imaging device 31 into the first trained model 2023. The frequency with which the image analysis module 2034 analyzes the image is periodic. The image analysis module 2034 may analyze the image, for example, weekly, every few days, daily, or every few hours. The image analysis may be performed, for example, at 9:00 p.m. every Sunday. Alternatively, the image analysis module 2034 may analyze the image immediately after the server 20 receives the image. In other words, the frequency with which the image analysis module 2034 analyzes the image can be set arbitrarily.
[0091] The first trained model 2023 outputs information related to the detected object.
[0092] In step S62, the control unit 203 determines whether or not the analyzed image contains an object that corresponds to the first learned model 2023. If a corresponding object is found, the control unit 203 shifts the processing to step S63. On the other hand, if a corresponding object is not found, the control unit 203 shifts the processing to step S64.
[0093] In step S63, the control unit 203 records information related to the object detected in step S62 in the detection log DB 2022. The recorded information includes, for example, the image capture device ID, the detection model, the detection category, the detection frame, and the image area.
[0094] In step S64, if any trained models that have not been used for detection remain, the control unit 203 repeats the processes of steps S61, S62, and S63 using the unused trained models. If any trained models that have not been used for detection remain, the control unit 203 ends the process. When the control unit 203 ends the process, the control unit 203 may notify the display 141 of the terminal device 10 that the control unit 203 has ended the process.
[0095] The user operates the terminal device 10 and accesses the server 20 after user authentication. The server 20 transmits information about the editing screen to the terminal device 10. The editing screen includes information in the detection log DB 2022 that has been updated since the previous access. The user edits the video on the editing screen.
[0096] FIG. 7 is a schematic diagram illustrating an editing screen displayed on the display 141 of the terminal device 10. The display area 71 is an area for displaying video in streaming format. Marks 721, 722, and 723 are borders surrounding objects detected in the video. The style of the borders may differ for each trained model or each detection category. The style of the borders may include, for example, color, thickness, shape, etc., or a combination thereof. Even if an object of only one detection model or only one detection category is detected in the video, a mark may be attached to the object. A numerical value 73 indicates the time point in the video being displayed. A group of icons 74 are icons for operating the video. The group of icons 74 is not limited to the three icons shown in FIG. 7.
[0097] The timeline 75 is a time axis for displaying the time associated with the detection marker 78 of each box 77 on a timeline. The timeline 75 is divided into sections at regular time points, for example. The playback marker 76 is a marker on the timeline 75 that indicates the time point of the video currently being played. The box 77 is a box that indicates the category of the detected object. The box 77 may also indicate information that can identify a trained model of the detected object. The detection marker 78 is a marker on the timeline 75 that indicates the time at which the object was detected. According to the example of FIG. 7 , the dog was detected between 8:00 and 9:00.
[0098] Icon 79 is an icon for editing the video. For example, icon 79 is an icon for dividing the video into certain time segments and then cutting out the video. For example, the user may determine the time of the video to be cut out by specifying detection marker 78. That is, image editing module 2035 cuts out the video for the time period corresponding to detection marker 78. Image editing module 2035 is not limited to matching the time period of the video to be cut out with detection marker 78, and may add a predetermined time width to detection marker 78. For example, image editing module 2035 may add a predetermined time width before the start point of detection marker 78, after the end point of detection marker 78, or both. Furthermore, the user may determine the time of the video to be cut out by, for example, specifying a time width on timeline 75. Specifically, for example, image editing module 2035 receives taps or drags at two points on timeline 75 from the user and cuts out the video for the received time width.
[0099] After editing the video, the image editing module 2035 saves the edited video file in, for example, the memory 25 or the storage 26.
[0100] 8 is a schematic diagram showing a portion of the editing screen displayed on the display 141 of the terminal device 10. As shown in FIG. 8, the user may determine the time period of the video to be cut out by, for example, specifying two or more detection markers. Specifically, as shown in editing case 81, the user may determine the time period of the video to be cut out by specifying two detection markers that are separated from each other on the timeline. In other words, the image editing module 2035 may cut out the video in the time periods corresponding to the detection markers specified by the user and combine the videos. The two specified detection markers may belong to the same category or different categories.
[0101] Specifically, as shown in editing case 82 and editing case 83, the user may determine the time of the video to be cut out by specifying two overlapping detection markers on the timeline. In editing case 82, the image editing module 2035 extracts the portion where the two detection markers specified by the user overlap. In editing case 83, the image editing module 2035 extracts the portion corresponding to at least one detection marker.
[0102] The icon 79 in Fig. 7 is not limited to the icon for cutting out shown in Fig. 7. For example, the icon 79 may be an icon for editing a moving image. Specifically, the icon 79 may be, for example, an icon for adjusting the brightness of a moving image or an icon for applying a predetermined effect to the moving image. In other words, the image editing module 2035 may accept an operation on the icon 79 from the user and edit the moving image.
[0103] The user may arbitrarily set rules for editing the video. Specifically, the user may set rules for automatic editing based on an object as a condition. For example, the user may set rules for automatically determining the time of the video to be cut out by specifying a child detection marker 78 and a toy detection marker 78. As a result, when a predetermined object is detected, a video including the object is automatically generated. Specifically, the user may set rules for automatic editing based on a time as a condition. For example, the user may set a rule for automatically increasing the brightness of a video shot at night. As a result, a video is automatically generated during a predetermined time period. In other words, the image editing module 2035 may automatically edit the video according to any video editing rules set by the user.
[0104] 7 is an icon for downloading the video edited by icon 79. In response to a user's operation on icon 710, the transmission control module 2032 transmits the edited video file from the memory 25, the storage 26, or the like to the terminal device 10. The terminal device 10 saves the downloaded video file.
[0105] As described above, in the above embodiment, the control unit 203 acquires a captured video. The control unit 203 analyzes the video using a plurality of pre-stored analytical models, each of which has been learned for a target object. The control unit 203 presents a detection marker 78 to the user, which indicates the timing at which the target object was detected by the analytical model, in association with the timeline 75 associated with the video. The control unit 203 generates an edited video from the captured video in response to a user instruction based on the detection marker 78. This allows the user to intuitively and easily grasp the situation at the scene where the video was captured via a user interface in which the appearance times of objects for each model are visually organized on the timeline 75. This also allows the user to intuitively and easily edit the video via a user interface for efficiently processing objects.
[0106] Therefore, the system according to this embodiment improves convenience when editing videos based on images captured by a surveillance camera or the like.
[0107] In the above embodiment, the control unit 203 receives a specification for the detection marker 78 in the generating step, and generates an edited video for the time period corresponding to the detection marker 78 or a time period with a predetermined range around the time period. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to not miss any signs before an object appears or any effects after the object leaves the scene.
[0108] In the above embodiment, the control unit 203 receives designations for multiple detection markers 78 in the generating step and generates an edited video of a combined time period by joining together time periods corresponding to the detection markers 78, or a combined time period by joining together time periods with a predetermined width therebetween. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to not miss any signs before an object appears or any effects after the object disappears.
[0109] In the above embodiment, if there is a time period in which the detection markers 78 detected by different analysis models overlap, the control unit 203 generates an edited video for the overlapping time period or for a time period that has a predetermined width within the time period. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to not miss any signs before an object appears or any effects after the object leaves the scene.
[0110] In the above embodiment, if there is a time period in which detection markers 78 based on different analysis models overlap, the control unit 203 generates an edited video for the time period corresponding to at least one of the detection markers 78 or a time period with a predetermined width around that time period in the generating step. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to not miss any signs before an object appears or any effects after the object leaves.
[0111] In the above embodiment, the control unit 203 receives a designation of a time period on the timeline 75 in the generating step and generates an edited video for the designated time period or a combined time period by joining together the designated time periods. This allows the user to edit the video more intuitively and easily via a user interface for efficiently processing objects. This also allows the user to edit the video without being limited by the detection markers 78 while still referring to the detection markers 78.
[0112] In the above embodiment, in the generating step, the control unit 203 receives instructions from the user specifying rules for generating an edited moving image based on the detected markers 78. This allows the user to edit the moving image more intuitively and simply via a user interface for efficiently processing objects. This also saves the user the trouble of inputting frequently occurring editing patterns.
[0113] In the above embodiment, the control unit 203 analyzes the video at an arbitrary frequency using the analysis model in the analysis step. As a result, the frequency of image analysis by the image analysis module 2034 can be set arbitrarily, so the device according to this embodiment does not necessarily require high performance.
[0114] In the above embodiment, in the presenting step, when objects detected using different analytical models are present in the same frame, the control unit 203 simultaneously presents the marks 721 to 723 representing the objects to the user. This allows the user to intuitively and easily grasp the situation of the scene photographed by the marks 721 to 723 attached to each object, even when multiple objects are detected in the same frame.
[0115] In the above embodiment, the control unit 203 presents the marks 721 to 723 in different forms for each analysis model in the presenting step, allowing the user to intuitively and easily grasp the situation of the site photographed by the marks 721 to 723, which have different forms for each analysis model.
[0116] <4 Other Embodiments> In the above embodiment, it has been described that the image capture device 31 transmits video to the server 20. However, the server 20 may be realized by a general information processing device such as a PC. In this case, for example, a storage medium storing video in the image capture device 31 is removed and connected to an input port of the information processing device, thereby inputting the information stored in the storage medium to the information processing device. The control unit of the information processing device performs processing similar to that of the image analysis module 2034 of the server 20, for example, analyzing images captured by the image capture device 31 and outputting information about objects at the site. In addition, the control unit of the information processing device performs processing similar to that of the image editing module 2035 of the server 20, for example, accepting editing instructions from a user and generating an edited video based on the accepted editing instructions.
[0117] In the above embodiment, it has been described that the edited video is downloaded from the memory 25, the storage 26, or the like to the terminal device 10 and stored therein. However, the user may view the edited video on the server 20 without downloading it.
[0118] In the above embodiment, it has been described that the image analysis module 2034 may limit the analysis range of an image. The number of analysis ranges for one image may be two or more. For example, the ranges analyzed for a video taken of a retail store may be two locations: the entrance and the exit. Furthermore, if objects are detected simultaneously in two or more analysis ranges, or if objects are detected within a certain time interval, the video for the time when the objects were detected may be extracted.
[0119] In the above embodiment, it has been described that a video of a time period equal to the detected marker plus a predetermined time width may be cut out. However, the video to be created is not limited to this. The image editing module 2035 may cut out a video of a time period consisting of only the predetermined time width.
[0120] In the above embodiment, the editing screen displays one video captured by one camera device 31 and one timeline corresponding to that video. However, what is displayed on the editing screen is not limited to these. In other words, when generating an edited video, the videos referenced are not limited to those captured by one camera device 31. For example, the image analysis module 2034 may analyze videos captured by two or more camera devices 31. Furthermore, the image editing module 2035 may generate an edited video based on videos captured by two or more camera devices 31. Specifically, for example, the image analysis module 2034 analyzes multiple videos captured by multiple camera devices 31 installed to capture the same site from different angles. The image editing module 2035 generates an edited video based on the analysis results.
[0121] 9 is a schematic diagram illustrating another example of an editing screen displayed on the display 141 of the terminal device 10. While FIG. 9 illustrates a case where one object is detected, multiple objects may be detected. In other words, multiple trained models may be used in the analysis. The editing screen in FIG. 9 includes, for example, multiple videos and one timeline collectively corresponding to the videos.
[0122] 9, a detection marker 78 indicates that an object has been detected in the video. As shown by the detection marker 78 in FIG. 9, the presentation module 2036 may display a detection marker for each video from each camera under the same object. By manipulating multiple detection markers under the same object, a user can extract videos based on, for example, the time at which the same object was detected in multiple videos.
[0123] As shown by the display area 71 in FIG. 9, multiple videos are displayed in divided display areas within the editing screen.
[0124] The image editing module 2035 may, for example, provide the user with options for the moving image to be edited. Specifically, for example, the image editing module 2035 may accept, from the user, a selection of moving images to be included in the edited moving image. For example, when the user selects one or more detection markers under the same object, the image editing module 2035 sets the moving images corresponding to the selected detection markers as the target for editing. The user may also select a moving image in which no detection marker appears under the same object.
[0125] If the user selects a single video, the image editing module 2035 edits the selected video in the same manner as if the editing screen displayed one video and one timeline corresponding to that video.
[0126] If the user selects multiple videos, the image editing module 2035 edits the selected videos, specifying how the multiple videos should be played, i.e., whether the selected videos should be played in sequence or in parallel.
[0127] When multiple videos are played in sequence, the order of playback is determined automatically or by user operation. The automatically determined order of playback may be determined, for example, based on time, the proportion of the screen occupied, etc. Specifically, the image editing module 2035 determines the order of playback taking into consideration, for example, the length of time indicated by the detection marker in the video, the average proportion of the area of the object that occupies the screen throughout the video, etc. Furthermore, when the user determines the order of playback, the image editing module 2035 may suggest the automatically determined order to the user as an aid in the decision.
[0128] When multiple videos are played in parallel, the multiple videos are played, for example, in a picture-in-picture format in which the secondary video is placed on top of the primary video in the video playback area. The primary-subordinate relationship between the videos is determined automatically or by user operation. The automatically determined primary-subordinate relationship may be determined based on, for example, time, the proportion of the screen occupied, etc. Specifically, the image editing module 2035 determines the primary-subordinate relationship by taking into consideration, for example, the length of time indicated by the detection marker in the video, the average proportion of the area of the object occupied on the screen throughout the video, etc. Furthermore, when the user determines the primary-subordinate relationship, the image editing module 2035 may suggest the automatically determined primary-subordinate relationship to the user as an aid in the determination.
[0129] Furthermore, when multiple videos are played in parallel, the multiple videos may be played in, for example, divided video playback areas. That is, the image editing module 2035 divides the video playback area so that multiple videos are played in the video playback area. The area division is determined automatically or in response to a user operation. The automatically determined division may be determined based on, for example, time, the proportion of the screen occupied, etc. Specifically, the image editing module 2035 determines the division taking into account, for example, the length of time indicated by the detection marker in the video, the average proportion of the area of the object occupied on the screen throughout the video, etc. Furthermore, when the user determines the division, the image editing module 2035 may suggest the automatically determined division to the user as an aid in the decision. Even if there are differences between the multiple videos in terms of time, the proportion of the screen occupied, etc., the image editing module 2035 may equally divide the video playback area in response to a user operation.
[0130] In the above embodiment, a detection marker related to a detected object is displayed on the timeline. However, information other than the detected object may also be displayed on the timeline. For example, the presentation module 2036 may display output information of a predetermined IoT device on the timeline. Specifically, the reception control module 2031 receives information measured by an IoT device attached to on-site equipment at a predetermined interval. The presentation module 2036 presents a detection marker based on the information transmitted from the IoT device on the timeline of the editing screen. For example, if the acquired information satisfies predetermined requirements, the presentation module 2036 presents the detection marker on the timeline of the editing screen. For example, if the reception control module 2031 receives information indicating a temperature above a certain level from a temperature sensor attached to equipment displayed on the image, the presentation module 2036 presents a detection marker indicating the time during which that temperature was maintained on the timeline. The presentation module 2036 may present the detection marker so that the detection marker indicates each temperature range in a different manner. For example, an orange detection marker may indicate a temperature between 30 degrees Celsius and 40 degrees Celsius. Furthermore, for example, when the reception control module 2031 receives information about acceleration above a certain level from an acceleration sensor attached to equipment shown on an image, the presentation module 2036 presents a detection marker on a timeline indicating the time during which that acceleration occurred. The presentation module 2036 may present the detection markers so that the detection markers indicate the acceleration ranges in different ways for each acceleration range. For example, a blue detection marker may indicate an acceleration between 1 G and 2 G.
[0131] 10 is a block diagram showing the basic hardware configuration of the computer 100. The computer 100 includes at least a processor 101, a main memory device 102, an auxiliary memory device 103, and a communication IF (interface) 109. These components are electrically connected to each other via a bus.
[0132] The processor 101 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.
[0133] The main memory device 102 is used to temporarily store programs and data to be processed by the programs, etc. For example, it is a volatile memory such as a DRAM (Dynamic Random Access Memory).
[0134] The auxiliary storage device 103 is a storage device for saving data and programs, such as a flash memory, a hard disk drive (HDD), a magneto-optical disk, a CD-ROM, a DVD-ROM, or a semiconductor memory.
[0135] The communication IF 109 is an interface for inputting and outputting signals for communicating with other computers via a network using a wired or wireless communication standard.
[0136] The network is composed of the Internet, a LAN, various mobile communication systems constructed by wireless base stations, etc. For example, the network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi (registered trademark)) that can connect to the Internet via a predetermined access point. In the case of a wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), and Bluetooth (registered trademark). In the case of a wired connection, the network also includes a network that is directly connected using a USB (Universal Serial Bus) cable, etc.
[0137] It should be noted that the computer 100 can be virtually realized by distributing all or part of each hardware configuration across multiple computers 100 and interconnecting them via a network. In this way, the concept of the computer 100 includes not only a computer 100 housed in a single housing or case, but also a virtualized computer system.
[0138] <Basic Functional Configuration of Computer 100> A description will be given of the functional configuration of the computer realized by the basic hardware configuration of the computer 100 shown in Fig. 10. The computer includes at least the functional units of a control unit, a storage unit, and a communication unit.
[0139] The functional units of the computer 100 can also be realized by distributing all or part of the functional units across multiple computers 100 interconnected via a network. The computer 100 is a concept that includes not only a single computer 100 but also a virtualized computer system.
[0140] The control unit is realized by the processor 101 reading various programs stored in the auxiliary storage device 103, loading them into the main storage device 102, and executing processing in accordance with the programs. The control unit can realize functional units that perform various types of information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.
[0141] The storage unit is realized by the main storage device 102 and the auxiliary storage device 103. The storage unit stores data, various programs, and various databases. Furthermore, the processor 101 can allocate a storage area corresponding to the storage unit in the main storage device 102 or the auxiliary storage device 103 in accordance with the programs. Furthermore, the control unit can cause the processor 101 to execute processes for adding, updating, and deleting data stored in the storage unit in accordance with the various programs.
[0142] The term "database" refers to a relational database, which manages data sets called tables in a tabular format structurally defined by rows and columns, by associating them with one another. In a database, a table is called a table, a column in a table, and a row in a table a record. In a relational database, relationships between tables can be set and associated. Typically, each table has a column set as a key to uniquely identify a record, but setting a key to a column is not required. The control unit can cause the processor 101 to add, delete, or update records in specific tables stored in the storage unit according to various programs.
[0143] The communication unit is realized by the communication IF 109. The communication unit realizes a function of communicating with other computers 100 via a network. The communication unit can receive information transmitted from other computers 100 and input the information to the control unit. The control unit can cause the processor 101 to execute information processing on the received information in accordance with various programs. The communication unit can also transmit information output from the control unit to other computers 100.
[0144] The functions performed by the components described herein may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs (Central Processing Units), conventional circuits, and / or combinations thereof, programmed to perform the described functions. Processors include transistors and other circuits and are considered circuitry or processing circuitry. A processor may also be a programmed processor that executes a program stored in memory. In this specification, a circuit, unit, or means is hardware that is programmed to perform or executes the described functions. The hardware may be any hardware disclosed herein or any hardware known to be programmed to perform or execute the described functions. When the hardware is a processor, which is considered a type of circuitry, the circuit, means, or unit is a combination of hardware and software used to configure the hardware and / or processor.
[0145] Although several embodiments of the present disclosure have been described above, these embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and modifications are intended to be included in the scope of the inventions and their equivalents as defined in the claims, as well as in the scope and spirit of the inventions.
[0146] <Supplementary Notes> The matters described in the above embodiments are supplemented below. (Supplementary Note 1) A program for operating a computer including a processor 29 and a memory 25, the program causing the processor to execute the steps of acquiring a captured video, analyzing the acquired video using a plurality of pre-stored analysis models, each of which has been trained to detect an object, presenting to a user detection markers 78 that indicate the timing at which the object was detected by the analysis models, in association with a timeline 75 associated with the video, and generating an edited video from the captured video in response to an instruction from the user based on the detection markers. (Supplementary Note 2) The program described in (Supplementary Note 1), wherein in the generating step, a specification for the detection marker is accepted, and an edited video is generated for a time period corresponding to the detection marker or for a time period within the time period by a predetermined range. (Supplementary Note 3) The program according to (Supplementary Note 1) or (Supplementary Note 2), wherein, in the generating step, a specification for a plurality of detection markers is accepted, and an edited video of a combined time period is generated by joining together time periods corresponding to the detection markers, or by joining together time periods with a predetermined width from the time period. (Supplementary Note 4) The program according to any of (Supplementary Note 1) to (Supplementary Note 3), wherein, in the generating step, if there is a time period in which detection markers based on different analysis models overlap, an edited video of the overlapping time period or a time period with a predetermined width from the time period is generated. (Supplementary Note 5) The program according to any of (Supplementary Note 1) to (Supplementary Note 4), wherein, in the generating step, if there is a time period in which detection markers based on different analysis models overlap, an edited video of at least a time period corresponding to any of the detection markers, or a time period with a predetermined width from the time period is generated. (Supplementary Note 6) The program according to any of (Supplementary Note 1) to (Supplementary Note 5), wherein, in the generating step, if there is a time period in which detection markers based on different analysis models overlap, an edited video of at least a time period corresponding to any of the detection markers, or a time period with a predetermined width from the time period.(Supplementary Note 7) The program according to any one of (Supplementary Note 1) to (Supplementary Note 6), wherein, in the generating step, a user provides rules as instructions for generating an edited video based on the detected markers. (Supplementary Note 8) The program according to any one of (Supplementary Note 1) to (Supplementary Note 7), wherein, in the analyzing step, an analytical model is used to analyze the video at any frequency. (Supplementary Note 9) The program according to any one of (Supplementary Note 1) to (Supplementary Note 8), wherein, in the presenting step, if objects detected by different analytical models are in the same frame, marks representing the objects are simultaneously presented to the user. (Supplementary Note 10) The program according to (Supplementary Note 9), wherein, in the presenting step, the marks have a different form for each analytical model. (Supplementary Note 11) A method executed by a computer comprising a processor and a memory, wherein the processor executes all of the steps executed in the invention according to any one of (Supplementary Note 1) to (Supplementary Note 10). (Supplementary Note 12) An information processing device comprising a control unit 203 and a storage unit 202, wherein the control unit executes all steps executed in the invention according to any one of (Supplementary Note 1) to (Supplementary Note 10). (Supplementary Note 13) A system comprising means for executing all steps executed in the invention according to any one of (Supplementary Note 1) to (Supplementary Note 10).
[0147] DESCRIPTION OF SYMBOLS 1...System 10...Terminal device 12...Communication IF 120...Communication unit 13...Input device 131...Touch-sensitive device 14...Output device 141...Display 15...Memory 150...Location information sensor 16...Storage 160...Camera 17...Audio processing unit 171...Microphone 172...Speaker 180...Memory unit 181...User information 19...Processor 190...Control unit 191...Operation acceptance unit 192...Transmission / reception unit 193...Presentation control unit 20...Server 201...Communication unit 202...Memory unit 2021...Photographing device DB 2022...Detection log DB 2023...First learned model 2024...Second learned model 2025...Third learned model 203...Control unit 2031...Reception control module 2032...Transmission control module 2034...Image analysis module 2035: Image editing module 2036: Presentation module 22: Communication IF 23: Input / output IF 25: Memory 26: Storage 29: Processor 31: Image capture device 80: Network 100: Server (computer) 101: Processor 102: Memory (main storage device) 103: Storage (auxiliary storage device) 109: Communication IF
Claims
1. A program for operating a computer having a processor and a memory, the program causing the processor to execute the following steps: acquiring a captured video; analyzing the captured video using a plurality of pre-stored analytical models, each of which has been trained to detect an object; presenting to a user detection markers that indicate the timing at which an object is detected by the analytical models, in association with a timeline associated with the video; and generating an edited video from the captured video in response to an instruction from the user based on the detection markers.
2. The program according to claim 1, wherein in the generating step, a specification for a detection marker is accepted, and the edited video is generated for a time period corresponding to the detection marker or for a time period with a predetermined width around that time period.
3. The program of claim 1, wherein in the generating step, specifications for a plurality of detection markers are accepted, and the edited video is generated for a combined time period by connecting time periods corresponding to the detection markers, or for a combined time period by connecting time periods with a predetermined width within the time period.
4. The program described in claim 1, wherein, in the generating step, if there is a time period in which detected markers obtained by different analysis models overlap, the edited video is generated for the overlapping time period or for a time period with a predetermined width therebetween.
5. The program described in claim 1, wherein, in the generating step, if there is a time period in which detection markers from different analysis models overlap, the edited video is generated for a time period corresponding to at least one of the detection markers, or for a time period with a predetermined width around that time period.
6. The program according to claim 1, wherein the generating step accepts specification of a time period on a timeline, and generates the edited video for the specified time period or a combined time period created by joining together these time periods.
7. The program according to claim 1, wherein in the generating step, the user provides, as the instruction, a rule for generating the edited video based on the detected markers.
8. The program according to claim 1, wherein, in the step of analyzing, the video is analyzed at an arbitrary frequency using the analysis model.
9. The program according to claim 1, wherein, in the step of presenting, when objects detected by different analysis models are in the same frame, marks representing the objects are presented to the user simultaneously.
10. The program according to claim 9, wherein, in the step of presenting, the marks are in different forms for each analysis model.
11. A method executed by a computer including a processor and a memory, wherein the processor executes all the steps executed in the invention according to any one of claims 1 to 10.
12. An information processing apparatus including a control unit and a storage unit, wherein the control unit executes all the steps executed in the invention according to any one of claims 1 to 10.
13. A system including means for executing all the steps executed in the invention according to any one of claims 1 to 10.
Citation Information
Patent Citations
Tracking support device, tracking support system and tracking support method
JP2017139701A
Information processing device, control method thereof, and program
JP2019102852A