Image processing system, image processing device, and image processing method
The video processing system addresses recognition accuracy issues by using multiple models and a switching mechanism that inputs historical data, ensuring seamless transitions and accurate event recognition.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2022-07-14
- Publication Date
- 2026-05-15
AI Technical Summary
Existing video recognition technologies struggle to adapt to changes in video environments, leading to reduced recognition accuracy when switching between different video analysis models.
A video processing system with multiple analysis models and a switching mechanism that inputs video data from a predetermined period before switching, ensuring continuous analysis and maintaining recognition accuracy.
The system maintains recognition accuracy by allowing both pre- and post-switch models to utilize past data, preventing recognition accuracy drops during model transitions.
Smart Images

Figure 0007859499000001 
Figure 0007859499000002 
Figure 0007859499000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a video processing system, a video processing apparatus, and a video processing method.
Background Art
[0002] Techniques for recognizing events in a video based on the video acquired via a network have been developed. For example, for video recognition that analyzes a video and recognizes events in the video, a recognition model using machine learning is utilized. The recognition model is also referred to as an analysis model or a recognition engine.
[0003] As related techniques, for example, Patent Documents 1 and 2 are known. Patent Document 1 describes a technique in which a first recognition engine and a second recognition engine recognize contexts respectively based on an input video. Further, Patent Document 1 also describes that a plurality of different types of recognition engines may be automatically selected every predetermined time.
[0004] Also, Patent Document 2 describes a technique for selecting a recognition engine for input data using a learning model learned by associating input data with an identifier of the recognition engine.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] As described above, related technologies such as those in Patent Documents 1 and 2 involve selecting a recognition model and analyzing the video using the selected model. However, in these related technologies, depending on the environment of the acquired video, it may not be possible to suitably recognize events within the video.
[0007] In view of these challenges, this disclosure aims to provide a video processing system, a video processing device, and a video processing method that are suitable for recognizing events within a video. [Means for solving the problem]
[0008] The video processing system according to this disclosure includes a first video analysis model for analyzing video corresponding to a first video recognition environment, a second video analysis model for analyzing video corresponding to a second video recognition environment, and a switching means for switching the video analysis model that analyzes the video input data from the first video analysis model to the second video analysis model in response to a change in the input video input data from the first video recognition environment to the second video recognition environment, wherein the switching means inputs video input data including data from a predetermined period prior to the switching timing to the second video analysis model in response to a change in the video input data from the first video recognition environment to the second video recognition environment.
[0009] The video processing device according to this disclosure includes a first video analysis model for analyzing video corresponding to a first video recognition environment, a second video analysis model for analyzing video corresponding to a second video recognition environment, and a switching means for switching the video analysis model that analyzes the video input data from the first video analysis model to the second video analysis model in response to a change in the input video input data from the first video recognition environment to the second video recognition environment, wherein the switching means inputs video input data including data from a predetermined period prior to the switching timing to the second video analysis model in response to a change in the video input data from the first video recognition environment to the second video recognition environment.
[0010] The video processing method relating to this disclosure involves switching the video analysis model that analyzes the video input data from a first video analysis model that analyzes video corresponding to the first video recognition environment to a second video analysis model that analyzes video corresponding to the second video recognition environment, in response to a change in the video input data from a first video analysis model that analyzes video corresponding to the first video recognition environment, and inputting video input data including data from a predetermined period prior to the switching timing into the second video analysis model, in response to a change in the video input data from the first video recognition environment to the second video recognition environment. [Effects of the Invention]
[0011] This disclosure provides an image processing system, an image processing device, and an image processing method that are suitable for recognizing events within an image. [Brief explanation of the drawing]
[0012] [Figure 1] This is a configuration diagram showing an overview of the video processing system according to the embodiment. [Figure 2] This is a configuration diagram showing an overview of the video processing device according to the embodiment. [Figure 3] This is a configuration diagram showing an overview of the video processing device according to the embodiment. [Figure 4] This is a flowchart outlining the video processing method according to the embodiment. [Figure 5] This is a diagram illustrating related image processing methods. [Figure 6] This is a diagram illustrating the image processing method according to an embodiment. [Figure 7] This is a configuration diagram showing the basic configuration of a remote monitoring system according to an embodiment. [Figure 8] This is a configuration diagram showing an example of the configuration of a remote monitoring system according to Embodiment 1. [Figure 9] This figure shows a specific example of the bitrate-recognition model table according to Embodiment 1. [Figure 10]It is a diagram showing a specific example of the recognition model-frame number table according to Embodiment 1. [Figure 11] It is a flowchart showing an operation example of the remote monitoring system according to Embodiment 1. [Figure 12] It is a configuration diagram showing a configuration example of the remote monitoring system according to Embodiment 2. [Figure 13] It is a configuration diagram showing a configuration example of the remote monitoring system according to Embodiment 3. [Figure 14] It is a diagram showing a specific example of the frame rate-recognition model table according to Embodiment 3. [Figure 15] It is a configuration diagram showing a configuration example of the remote monitoring system according to Embodiment 4. [Figure 16] It is a diagram for explaining an operation example of the remote monitoring system according to Embodiment 4. [Figure 17] It is a diagram showing a specific example of the packet loss-recognition model table according to Embodiment 5. [Figure 18] It is a configuration diagram showing a configuration example of the remote monitoring system according to Embodiment 6. [Figure 19] It is a diagram showing a specific example of the scene-recognition model table according to Embodiment 6. [Figure 20] It is a configuration diagram showing a configuration example of the remote monitoring system according to Embodiment 7. [Figure 21] It is a diagram showing a specific example of the object size-recognition model table according to Embodiment 7. [Figure 22] It is a configuration diagram showing a configuration example of the remote monitoring system according to Embodiment 8. [Figure 23] It is a diagram showing a specific example of the operation speed-recognition model table according to Embodiment 8. [Figure 24] It is a configuration diagram showing a configuration example of the remote monitoring system according to Embodiment 9. [Figure 25] It is a diagram showing a specific example of the shooting state-recognition model table according to Embodiment 9. [Figure 26] It is a configuration diagram showing a configuration example of the remote monitoring system according to Embodiment 10. [Figure 27] This figure shows a specific example of the computation amount-recognition model table according to Embodiment 10. [Figure 28] This is a configuration diagram showing an example of the configuration of a remote monitoring system according to Embodiment 11. [Figure 29] This figure shows a specific example of a transmission bandwidth-recognition model table according to Embodiment 11. [Figure 30] This is a configuration diagram showing an overview of the computer hardware according to the embodiment. [Modes for carrying out the invention]
[0013] The embodiments will be described below with reference to the drawings. In each drawing, the same elements are denoted by the same reference numerals, and redundant explanations will be omitted where necessary.
[0014] (Summary of the embodiment) First, an overview of the embodiment will be described. Figure 1 shows the general configuration of the video processing system 10 according to the embodiment. The video processing system 10 can be applied, for example, to a remote monitoring system that collects video via a network and analyzes the video.
[0015] As shown in Figure 1, the video processing system 10 includes recognition models M1 and M2 and a switching unit 11. Recognition model M1 is a first video analysis model that analyzes video corresponding to a first video recognition environment. Recognition model M2 is a second video analysis model that analyzes video corresponding to a second video recognition environment. Recognition models M1 and M2 recognize, for example, a person's face, a vehicle, equipment, etc., depending on the input video. In addition, recognition models M1 and M2 may recognize, for example, a person's actions, the driving status of a vehicle, the state of an object, etc. Note that the recognition targets recognized by recognition models M1 and M2 are not limited to these examples. The video processing system 10 is not limited to two recognition models, but may include three or more recognition models. For example, recognition model M1 may be generated by training it with video training data corresponding to a first video recognition environment, and recognition model M2 may be generated by training it with video training data corresponding to a second video recognition environment. Alternatively, the created recognition models may be acquired and evaluated. For example, the recognition accuracy of multiple created recognition models may be evaluated using video corresponding to the first video recognition environment, and the recognition model with the highest accuracy may be determined to be recognition model M1 used in the first video recognition environment. Similarly, the recognition accuracy of multiple created recognition models may be evaluated using video corresponding to the second video recognition environment, and the recognition model with the highest accuracy may be determined to be recognition model M2 used in the second video recognition environment.
[0016] The video recognition environment is the video environment that the recognition model analyzes and recognizes, and may indicate the quality of the video, or it may indicate the environment including the objects that appear in the video. Note that "analysis and recognition" means that either analysis or recognition is performed. The video recognition environment may also include, for example, video parameters such as bitrate and frame rate that indicate the quality of the video, the communication quality of the video received via the network, the scene in which the video was taken, the size of the objects included in the video, the speed of movement of the objects included in the video, and the shooting conditions under which the video was taken. A scene could be, for example, the progress of a construction project, the work content and work location of workers, etc.
[0017] The switching unit 11 switches the recognition model, i.e., the video analysis model, that analyzes the video input data in response to a change from a first video recognition environment to a second video recognition environment in the input video data. The video input data is video data that recognition models M1 or M2 analyze and perform recognition processing on, and includes, for example, recognition targets such as human faces, vehicles, and equipment. When video input data is input to recognition models M1 and M2, recognition models M1 and M2 may perform analysis and recognition processing. In response to a change from a first video recognition environment to a second video recognition environment in the video input data, the switching unit 11 inputs video input data including data from a predetermined period before the switching timing to the target recognition model M2. That is, the switching unit 11 inputs data from a predetermined period before the switching timing up to the switching timing to recognition model M2, and further inputs data from the switching timing onward to recognition model M2. The same applies when switching from recognition model M2 to recognition model M1.
[0018] The switching unit 11 may input video input data, including the number of frames used by the target recognition model M2 for image recognition, to the target recognition model M2 as video input data that includes data from a predetermined period prior to the switching timing. Alternatively, the switching unit 11 may input video input data that includes data from a predetermined period prior to the switching timing to both the source recognition model M1 and the target recognition model M2. In other words, the switching unit 11 may input data from a predetermined period prior to the switching timing up to the switching timing to recognition models M1 and M2.
[0019] The video processing system 10 may consist of one device or multiple devices. Figure 2 illustrates the configuration of a video processing device 20 according to an embodiment. As shown in Figure 2, the video processing device 20 may include the recognition models M1 and M2 and the switching unit 11 shown in Figure 1. Furthermore, part or all of the video processing system 10 may be located at the edge or in the cloud. For example, the recognition models M1 and M2 and the switching unit 11 may be located on a cloud server. In addition, each function may be distributed across the cloud. Figure 3 illustrates a configuration in which the functions of the video processing system 10 are located on multiple video processing devices. In the example in Figure 3, the video processing device 21 includes the switching unit 11, and the video processing device 22 includes the recognition models M1 and M2. Note that the configuration in Figure 3 is just one example and is not limited to this configuration.
[0020] Furthermore, recognition models M1 and M2 may be placed at the same location or at different locations. For example, recognition model M1 may be placed at one of the edge and cloud locations, and recognition model M2 may be placed at the other of the edge and cloud locations.
[0021] Figure 4 shows an image processing method according to an embodiment. For example, the image processing method according to the embodiment is performed by the image processing system 10 in Figure 1 or the image processing devices 20-22 in Figure 2 or Figure 3. As shown in Figure 4, in response to the change from a first image recognition environment to a second image recognition environment in the input video data (S11), the recognition model that analyzes the video input data, i.e., the video analysis model, is switched from the recognition model M1 that analyzes the video corresponding to the first image recognition environment to the recognition model M2 that analyzes the video corresponding to the second image recognition environment (S12). Also, in response to the change from the first image recognition environment to the second image recognition environment in the video input data (S11), video input data including data from a predetermined period prior to the switching timing is input to the recognition model M2 (S13).
[0022] Here, we will describe the challenges in related technologies prior to the application of the embodiment. Specifically, we will examine a video processing method in which video is transmitted from a terminal to a server and the server switches the recognition model, using related technologies such as Patent Documents 1 and 2.
[0023] Figure 5 illustrates the operation of selecting and switching between recognition models M1 and M2 in Figure 1 in a related video processing method. For example, recognition models M1 and M2 are models that learn and analyze video with different bitrates or compression ratios. In this example, the video to be captured and analyzed includes frames F1 to F8... arranged in chronological order, and the system switches from recognition model M1 to recognition model M2 at the timing of frame F8. Here, as an example, compressed and decompressed video is input to the recognition models, but the system is not limited to this configuration as long as it is possible to input video that can be analyzed and recognized into each recognition model. For example, a video processing system that performs the video processing method in Figure 5 may further include, in addition to the configuration in Figure 1, a shooting unit for capturing video, a compression unit for compressing video, and a decompression unit for decompressing compressed video. For example, a video processing system that performs the video processing method in Figure 5 does not necessarily have to include a compression unit and a decompression unit.
[0024] As shown in Figure 5, in the related video processing method, the shooting unit shoots video (S901), and the compression unit compresses the shot video (S902). Next, the compression unit sends the compressed video to the restoration unit, which restores the received compressed video to the original video (S903). Next, the switching unit selects the recognition model M1 and inputs frames F1 to F7 to the recognition model M1 before switching (S904). The recognition model M1 before switching performs video recognition using the input frames F1 to F7.
[0025] Next, at the switching timing, the switching unit switches the recognition model from M1 to M2 and inputs frames from frame F8 onwards to the switched recognition model M2 (S905). The switched recognition model M2 performs image recognition using the input frames from frame F8 onwards.
[0026] The inventors investigated the recognition accuracy when switching recognition models in the related video processing method, as shown in Figure 5, and found the following problems. Specifically, in a video processing method that analyzes by switching between multiple models, if the recognition model uses analysis information from past frames, sufficient analysis accuracy may not be obtained even after switching the recognition model. In other words, in a video recognition model that recognizes events using video, changing the recognition model into which the video is input, as in the related video processing method, may reduce the recognition accuracy of the target recognition model.
[0027] The recognition model is a video recognition engine that uses machine learning. For example, during training, it is a model that learns the movements of the person to be recognized based on time-series video data. The recognition model extracts the temporal changes in each frame of the video data and learns the movements of the person. Therefore, even during recognition, it is assumed that time-series video data is input to the recognition model, and it is necessary to input video with a sufficient number of frames to extract the temporal changes in each frame of the video data, even during recognition.
[0028] However, in the example in Figure 5, when switching from recognition model M1 to recognition model M2, the input to the new recognition model M2 starts from frame F8 after the switch. Therefore, recognition model M2 only receives video data from frame F8 onwards. Consequently, recognition model M2 does not receive past data prior to frame F8, and immediately after the switch, that is, at the moment of the switch, recognition model M2 cannot analyze the time-series data. For this reason, immediately after the switch, the recognition accuracy, i.e., the analysis accuracy, of the new recognition model M2 may decrease, or it may not be possible to obtain a recognition result at all. Recognition model M2 may not be able to correctly analyze using past data, which could lead to misrecognition of the object to be recognized in the video, and it may also be unable to output a recognition result.
[0029] Specific examples of such problems include cases where, even if only video footage of a person opening a vehicle door is input to a recognition model, it cannot determine whether the person is getting into or out of the vehicle; cases where only video footage of a person walking is input to a recognition model, but it cannot determine whether the person is walking forward or backward; and cases where only video footage of a person or machine holding an object is input to a recognition model, but it cannot determine whether the person or machine is lifting or lowering the object.
[0030] Therefore, in this embodiment, as shown in Figures 1 to 4, when switching recognition models, the data from the previous model is input to the target recognition model. Figure 6 illustrates the operation when switching recognition models at the same timing as in Figure 5 in the video processing method according to this embodiment. In this example as well, similar to Figure 5, for example, recognition models M1 and M2 are models that learn and analyze videos with different bitrates or compression ratios. As an example, compressed and decompressed videos are input to the recognition models, but the configuration is not limited to this as long as each recognition model can be input with videos that can be analyzed and recognized. For example, the video processing system that executes the video processing method of Figure 6 may further include, in addition to the configuration of Figure 1, a shooting unit that captures video, a compression unit that compresses video, and a decompression unit that decompresses the compressed video. For example, the video processing system that executes the video processing method of Figure 6 does not have to include a compression unit and a decompression unit.
[0031] As shown in Figure 6, in the video processing method according to the embodiment, similar to Figure 5, the shooting unit shoots video (S101), the compression unit compresses the shot video (S102), and the decoding unit restores the compressed video to the original video (S103). Next, the switching unit selects the recognition model M1 and inputs frames F1 to F7 to the recognition model M1 before switching (S104). The recognition model M1 before switching performs video recognition using the input frames F1 to F7.
[0032] Next, in this embodiment, the switching unit inputs frames F5 to F7, which are before the switching timing, to the recognition model M1 before the switching and the recognition model M2 after the switching (S105). Then, at the switching timing, the switching unit switches the recognition model from M1 to M2 and inputs frames F8 and later to the recognition model M2 after the switching (S106). As a result, the recognition model M2 after the switching performs image recognition using frames F5 and later, which are input before the switching timing.
[0033] In this embodiment, a frame from slightly before the model switch is input to both the pre- and post-switch recognition models. This allows the post-switch recognition model to perform image recognition using past data immediately after the switch, thus preventing a decrease in recognition accuracy or interruption of analysis. Furthermore, the target recognition model only needs to be input a number of frames sufficient to extract the temporal changes in each frame of the image data. Therefore, since only a few frames are needed to input to both recognition models, it is possible to suppress a decrease in recognition accuracy while maintaining approximately the same processing load for both recognition models compared to related technologies. In other words, while continuously inputting data to both recognition models increases the processing load, this increase can be suppressed by inputting only a predetermined number of frames from the time of the switch to both recognition models.
[0034] (Basic configuration of a remote monitoring system) Next, we will describe a remote monitoring system, which is an example of a system to which the embodiment is applied. Figure 7 illustrates the basic configuration of the remote monitoring system 1. The remote monitoring system 1 is a system that monitors an area captured by a camera using the captured video. In this embodiment, it will be described as a system that remotely monitors the work of workers at a site. For example, the site may be a work site such as a construction site, a public square where people gather, a school, or any area where people or machines are operating. In this embodiment, the work will be described as construction work or civil engineering work, but is not limited to these. Since video includes multiple images (also called frames) in a time series, video and images are interchangeable. That is, the remote monitoring system can be said to be a video processing system that processes video, and also an image processing system that processes images.
[0035] As shown in Figure 7, the remote monitoring system 1 includes multiple terminals 100, a central server 200, a base station 300, and an MEC 400. The terminals 100, base station 300, and MEC 400 are located on the field side, while the central server 200 is located on the central side. For example, the central server 200 is located in a data center or similar location far from the field. The field side is also referred to as the system edge side, and the central side is also referred to as the cloud side.
[0036] Terminal 100 and base station 300 are connected via network NW1, enabling communication. Network NW1 is a wireless network such as 4G, local 5G / 5G, LTE (Long Term Evolution), or Wi-Fi. Note that network NW1 is not limited to a wireless network but may also be a wired network. Base station 300 and central server 200 are connected via network NW2, enabling communication. Network NW2 includes, for example, core networks such as 5GC (5th Generation Core network) or EPC (Evolved Packet Core), or the Internet. Note that network NW2 is not limited to a wired network but may also be a wireless network. It can also be said that terminal 100 and central server 200 are connected via base station 300, enabling communication. Base station 300 and MEC 400 are connected via any communication method, but base station 300 and MEC 400 may be a single device.
[0037] Terminal 100 is a terminal device connected to the network NW1 and also serves as a video acquisition device for acquiring on-site video. Terminal 100 acquires video captured by camera 101 installed on-site and transmits the acquired video to the center server 200 via base station 300. Camera 101 may be located outside or inside terminal 100.
[0038] Terminal 100 compresses the video from camera 101 to a predetermined bitrate and transmits the compressed video. Terminal 100 has a compression efficiency optimization function 102 that optimizes the compression efficiency. The compression efficiency optimization function 102 performs RoI (Region of Interest; also called the gaze area) control to control the image quality of the RoI. The compression efficiency optimization function 102 reduces the bitrate by maintaining the image quality of the ROI, which includes people and objects, while lowering the image quality of the surrounding area.
[0039] Base station 300 is a base station device for network NW1 and also a relay device that relays communication between terminal 100 and central server 200. For example, base station 300 may be a local 5G base station, a 5G gNB (next generation node B), an LTE eNB (evolved node B), a wireless LAN access point, etc., but other relay devices may also be used.
[0040] The MEC (Multi-access Edge Computing) 400 is an edge processing unit located at the edge of the system. The MEC 400 is an edge server that controls terminals 100 and has a compressed bitrate control function 401 that controls the bitrate of the terminals. The compressed bitrate control function 401 controls the bitrate of terminals 100 through adaptive video distribution control and QoE (quality of experience) control. Adaptive video distribution control controls the bitrate of the video to be distributed according to the network conditions. For example, the compressed bitrate control function 401 predicts the recognition accuracy that can be obtained while suppressing the bitrate according to the communication environment of networks NW1 and NW2, and allocates the bitrate to the camera 101 of each terminal 100 in order to improve the recognition accuracy.
[0041] The center server 200 is a server located on the central side of the system. The center server 200 may be one or more physical servers, or it may be a cloud server or other virtualized server built on the cloud. The center server 200 is a monitoring device that monitors on-site work by analyzing on-site camera footage. The center server 200 is also a video analysis device that analyzes video transmitted from terminal 100.
[0042] The center server 200 has a video recognition function 201, an alert generation function 202, a GUI drawing function 203, and a screen display function 204. The video recognition function 201 recognizes the type of work performed by a worker, i.e., the type of human action, by inputting video transmitted from the terminal 100 into a video recognition AI (Artificial Intelligence) engine. The video recognition function 201 may include multiple recognition models, i.e., video analysis models, that analyze video corresponding to different video recognition environments. Furthermore, the center server 200 may be equipped with a switching unit that switches the recognition model according to changes in the video recognition environment. The alert generation function 202 generates alerts according to the recognized work. The GUI drawing function 203 displays a GUI (Graphical User Interface) on the display device screen. The screen display function 204 displays video from the terminal 100, recognition results, alerts, etc. on the GUI. Note that any of the functions may be omitted or any of the functions may be provided as needed. For example, the center server 200 may not have the alert generation function 202, the GUI drawing function 203, or the screen display function 204.
[0043] (Embodiment 1) Next, Embodiment 1 will be described. In this embodiment, as a change in the video recognition environment, an example of switching the recognition model in accordance with a change in the video bitrate will be described.
[0044] First, the configuration of the remote monitoring system according to this embodiment will be described. The basic configuration of the remote monitoring system 1 according to this embodiment is as shown in Figure 7. Figure 8 shows an example of the configuration of the remote monitoring system 1 according to this embodiment. Note that the configuration of each device is just an example, and other configurations are also acceptable as long as the operation according to this embodiment, which will be described later, is possible. For example, some functions of terminal 100 may be placed in the center server 200 or other devices, or some functions of the center server 200 may be placed in terminal 100 or other devices. In addition, the functions of MEC400, including the compressed bitrate control function, may be placed in the center server 200, etc.
[0045] As shown in Figure 8, the terminal 100 according to this embodiment includes a video acquisition unit 110, an encoder 120, and a terminal communication unit 130.
[0046] The video acquisition unit 110 acquires the video captured by the camera 101. The video captured by the camera is also referred to as the input video. For example, the input video may include people, such as workers, performing tasks on site. The video acquisition unit 110 is also an image acquisition unit that acquires multiple images in a time series, i.e., frames.
[0047] The encoder 120 encodes the acquired input video. The encoder 120 is an encoding unit that encodes the input video. The encoder 120 is also a compression unit that compresses the input video using a predetermined encoding scheme. The encoder 120 encodes using a video encoding scheme such as H.264 or H.265. The encoder 120 may also detect ROIs that include people and encode the input video so that the detected ROIs have higher image quality than other areas. An ROI identification unit may be provided between the video acquisition unit 110 and the encoder 120. The ROI identification unit detects objects in the acquired video and identifies regions such as ROIs. The encoder 120 may encode the input video so that the ROI identified by the ROI identification unit is of higher quality than other regions. Alternatively, the encoder 120 may encode the input video so that the region specified by the ROI identification unit is of lower quality than other regions. When detecting or identifying an ROI, the ROI identification unit or encoder 120 may maintain information corresponding to objects that may appear in the video and their priority, and identify regions such as ROIs according to this priority correspondence information.
[0048] The encoder 120 encodes the input video at a predetermined bitrate. The encoder 120 may encode the input video to the bitrate, frame rate, etc., assigned by the compression bitrate control function 401 of the MEC 400. Alternatively, the encoder 120 may determine the bitrate, frame rate, etc., based on the communication quality between the terminal 100 and the center server 200. Communication quality may be, for example, communication speed, but may also be other indicators such as transmission delay or error rate. The terminal 100 may be equipped with a communication quality measurement unit to measure communication quality. For example, the communication quality measurement unit determines the bitrate of the video to be transmitted from the terminal 100 to the center server 200 according to the communication speed. The communication speed may be measured based on the amount of data received by the base station 300 or the center server 200, and the communication quality measurement unit may acquire the communication speed measured from the base station 300 or the center server 200. Alternatively, the communication quality measurement unit may estimate the communication speed based on the amount of data per unit time transmitted from the terminal communication unit 130.
[0049] The terminal communication unit 130 transmits the encoded data (compressed data) encoded by the encoder 120 to the center server 200 via the base station 300. The terminal communication unit 130 is a transmission unit that transmits the acquired input video via the network. The terminal communication unit 130 is an interface that can communicate with the base station 300, and is a wireless interface such as 4G, local 5G / 5G, LTE, or Wi-Fi, but may also be a wireless or wired interface of any other communication method.
[0050] Furthermore, as shown in Figure 8, the center server 200 according to this embodiment includes recognition models M11 and M12, a center communication unit 210, a decoder 220, a prediction unit 230, a determination unit 240, a switching unit 250, and a storage unit 260.
[0051] Recognition models M11 and M12 perform video recognition processing on the input video. In this example, video recognition processing is performed on the received video received from the terminal and decoded. The video recognition processing is, for example, action recognition processing that recognizes the actions of people in the video, but other types of recognition processing may also be used. Recognition models M11 and M12 detect objects from the received video, recognize the actions of the detected objects, and output the results of the action recognition.
[0052] Recognition models M11 and M12 are video recognition engines that utilize machine learning, such as deep learning. By machine learning the features and behavioral labels of people performing tasks in video footage, they can recognize the actions of people in the video. For example, recognition models M11 and M12 are learning models that can learn and predict based on time-series video data, and can be CNNs (Convolutional Neural Networks), RNNs (Recurrent Neural Networks), or other neural networks.
[0053] Recognition models M11 and M12 are models trained on video from different video recognition environments as training data, and are learning models for analyzing video from different video recognition environments. Recognition model M11 is trained on video from a first video recognition environment, and recognition model M12 is trained on video from a second video recognition environment. Recognition models M11 and M12 can accurately analyze video from the video recognition environments they have each trained on. Therefore, if the video recognition environment of the received video is the first video recognition environment... place In addition, the received video is analyzed by the recognition model M11, and the video recognition environment of the received video is the second video recognition environment. place In addition, by analyzing the received video with the recognition model M11, the video can be analyzed with high accuracy.
[0054] The video recognition environment consists of video parameters related to video quality, such as bitrate and frame rate. It is not limited to bitrate and frame rate; it may also include compression ratio, image resolution, etc. In this embodiment, an example of bitrate will be described. Recognition model M11 learns video within a first bitrate range, and recognition model M12 learns video within a second bitrate range. Note that it is not limited to a first bitrate range and a second bitrate range; it may also be a first bitrate and a second bitrate. For example, the first bitrate range is a higher bitrate range than the second bitrate range, and recognition model M11 is a high-bitrate model, while recognition model M12 is a low-bitrate model, but it is not limited to this. Note that the first bitrate range and the second bitrate range may partially overlap.
[0055] The central communication unit 210 receives encoded data transmitted from the terminal 100 via the base station 300. The central communication unit 210 is a receiving unit that receives input video acquired by the terminal 100 via the network. The central communication unit 210 is an interface that can communicate with the internet or a core network, for example, a wired interface for IP communication, but it may also be a wired or wireless interface of any other communication method.
[0056] Decoder 220 decodes the encoded data received from terminal 100. Decoder 220 is a decoding unit that decodes the encoded data. Decoder 220 is also a restoration unit that restores the encoded data, i.e., compressed data, using a predetermined encoding scheme. Decoder 220 corresponds to the encoding scheme of terminal 100 and decodes using a video encoding scheme such as H.264 or H.265. Decodes according to the compression ratio and bitrate of each region and generates the decoded video. The decoded video will hereafter be referred to as the received video.
[0057] The prediction unit 230 predicts changes in the video recognition environment in the decoded received video. The prediction unit 230 extracts information about the video recognition environment from the received video and predicts changes in the video recognition environment by monitoring the extracted information. For example, the prediction unit 230 predicts changes in the bitrate extracted from the received video.
[0058] The determination unit 240 determines a recognition model to analyze the received video according to the video recognition environment of the received video, and determines the timing for switching the recognition model according to the predicted change in the video recognition environment. For example, the determination unit 240 determines a recognition model to analyze the received video according to the bitrate extracted from the received video. The determination unit 240 also determines the target recognition model and switching timing based on the bitrate change predicted by the prediction unit 230. Furthermore, based on the recognition model switching timing, the determination unit 240 determines a pre-input timing in which video data is input to the target recognition model in advance when switching. The pre-input timing is the timing in which video data is input to the target recognition model a predetermined period before the switching timing.
[0059] For example, the pre-input timing may be determined based on the number of pre-input frames that are input to the recognition model in advance. The number of pre-input frames is the number of frames that the target recognition model uses to perform image recognition. The number of pre-input frames is also the number of frames that are input to both recognition models when switching. Since the number of pre-input frames differs depending on the target recognition model, it is set in advance for each recognition model. For example, the number of pre-input frames may be changed according to the required recognition accuracy. In addition, a predetermined period corresponding to the number of pre-input frames may be associated with each recognition model.
[0060] The switching unit 250 switches between recognition models M11 and M12, which analyze the decoded received video. Based on the recognition model determined by the determination unit 240, the switching unit 250 selects a recognition model and inputs the received video to the selected recognition model. Based on the determined target recognition model and switching timing, the switching unit 250 switches the recognition model to which the received video is input. Based on the determined pre-input timing, the switching unit 250 inputs the video to the target recognition model before the switching timing. Between the pre-input timing and the switching timing, the switching unit 250 inputs the video to both the pre-switch recognition model and the post-switch recognition model.
[0061] The memory unit 260 stores data necessary for processing by the center server 200. The memory unit 260 stores a video recognition environment-recognition model table that associates video recognition environments with recognition models. Figure 9 shows a specific example of a bitrate range-recognition model table, which associates bitrate ranges with recognition models, as an example of a video recognition environment-recognition model table. The bitrate range-recognition model table allows the selection of a recognition model to analyze video according to the video's bitrate. In this example, bitrate range R1 is associated with recognition model M11, and bitrate range R2 is associated with recognition model M12. Bitrate ranges R1 and R2 correspond to the video bitrate ranges learned by each recognition model. For example, bitrate range R1 is a high bitrate range higher than bitrate range R2, and bitrate range R2 is a low bitrate range lower than bitrate range R1.
[0062] Furthermore, the memory unit 260 stores a recognition model-frame count table that associates a recognition model with the number of pre-input frames. Figure 10 shows a specific example of the recognition model-frame count table. The recognition model-frame count table allows the number of pre-input frames to be determined according to the target recognition model. In this example, frame count N1 is associated with recognition model M11, and frame count N2 is associated with recognition model M12. In addition to the number of frames, a pre-input time, which is a predetermined period corresponding to the number of frames to be input in advance, may be associated with the recognition model, and the pre-input timing may be determined from the pre-input time according to the target recognition model.
[0063] Next, the operation of the remote monitoring system according to this embodiment will be described. Figure 11 shows an example of the operation of the remote monitoring system 1 according to this embodiment. For example, the explanation assumes that terminal 100 executes S111 to S113 and the center server 200 executes S114 to S122, but this is not limited to this, and any device may execute each process. Some functions of the central server 200 may be assigned to other devices, and these other devices may perform those functions. For example, terminal 100 and MEC 400 may be equipped with a prediction unit 230, a decision unit 240, a switching unit 250, and a storage unit 260. Terminal 100 and MEC 400 may predict changes in the video recognition environment based on acquired video and changes in communication quality, determine the recognition model and switching timing by referring to the information in the storage unit, and notify the central server 200 of the switching timing instruction. Note that, not limited to this embodiment, terminal 100 and MEC 400 may similarly be equipped with a prediction unit 230, a decision unit 240, a switching unit 250, and a storage unit 260 in other embodiments as well.
[0064] As shown in Figure 11, terminal 100 acquires video from camera 101 (S111). Camera 101 generates video of the site, and video acquisition unit 110 acquires the video output from camera 101 (input video). For example, the image of the input video includes people working at the site and objects used in the work.
[0065] Next, terminal 100 encodes the acquired input video (S112). Encoder 120 encodes the input video using a predetermined video encoding method. For example, encoder 120 may encode the input video to a bitrate assigned by the compression bitrate control function 401 of MEC 400, or it may encode it at a bitrate corresponding to the communication quality between terminal 100 and the center server 200.
[0066] Next, terminal 100 sends the encoded data to the central server 200 (S113), and the central server 200 receives the encoded data (S114). The terminal communication unit 130 transmits encoded data, which is the input video, to the base station 300. The base station 300 forwards the received encoded data to the central server 200 via the core network or the internet. The central communication unit 210 receives the forwarded encoded data from the base station 300.
[0067] Next, the center server 200 decodes the received encoded data (S115). The decoder 220 decodes the encoded data according to the compression ratio and bitrate of each region and generates the decoded video, i.e., the received video.
[0068] Furthermore, the center server 200 predicts changes in the bitrate of the received video (S116). As an example of a video recognition environment, the prediction unit 230 monitors the bitrate of the received video and predicts changes in the bitrate. For example, the prediction unit 230 measures the amount of data per unit time in the encoded data received by the center communication unit 210 and obtains the bitrate. Alternatively, the terminal 100 may send a packet containing encoded data and bitrate, and the prediction unit 230 may obtain the bitrate from the received packet. Based on the history of past bitrates obtained periodically, the prediction unit 230 extracts the trend of bitrate transitions and predicts subsequent changes in bitrate.
[0069] Next, the center server 200 determines the switching timing (S117). The determination unit 240 determines the target recognition model and switching timing according to the predicted change in bitrate. The determination unit 240 refers to the bitrate range-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted bitrate. In the example bitrate range-recognition model table in Figure 9, if it is predicted that the bitrate of the received video will change from bitrate range R1 to bitrate range R2, it is decided to switch the recognition model from M11 to M12, and the timing at which the bitrate changes from bitrate range R1 to bitrate range R2 is determined as the switching timing. For example, the predicted bitrate is compared with the center of bitrate range R1 and the center of bitrate range R2, and the timing at which the predicted bitrate changes from a state close to the center of bitrate range R1 to a state close to the center of bitrate range R2 is set as the switching timing.
[0070] Next, the center server 200 determines the pre-input timing (S118). Based on the determined switching timing of the recognition model, the determination unit 240 determines the pre-input timing for inputting video data to the target recognition model in advance during the switching. The determination unit 240 refers to the recognition model-frame count table in the storage unit 260 and determines the number of pre-input frames corresponding to the target recognition model. In the example of the recognition model-frame count table in Figure 10, if the target recognition model is M12, the number of pre-input frames is determined to be N2. Furthermore, the pre-input time corresponding to the number of pre-input frames N2 is calculated based on the frame rate, and the pre-input timing is determined by subtracting the pre-input time from the switching timing.
[0071] Next, the center server 200 switches the input of the received video to the recognition model (S119). The switching unit 250 selects a recognition model according to the determined pre-input timing and switching timing, and inputs the decoded received video to the selected recognition model (S120~S122).
[0072] Specifically, if the current time is before the pre-input timing, the switching unit 250 inputs the received video to the recognition model before the switch (S120). For example, the switching unit 250 inputs the received video (frame) only to the recognition model M11 before the switch. The recognition model M11 performs video recognition using the input received video.
[0073] Furthermore, if the current time falls between the pre-input timing and the switching timing, the switching unit 250 inputs the received video to the recognition models before and after the switching (S121). For example, the switching unit 250 inputs frames of the received video to both the recognition model M11 before the switching and the recognition model M12 after the switching. Recognition model M11 performs video recognition using the received video input from S120 and outputs the recognition result. Recognition model M12 starts video recognition processing or makes video recognition processing possible using the received video input from S121.
[0074] Furthermore, if the current time is after the switching timing, the switching unit 250 inputs the received video to the recognition model after the switch (S122). For example, the switching unit 250 inputs the frames of the received video only to the recognition model M12 after the switch. The recognition model M12 performs video recognition using the received video input from S121 and outputs the recognition result. The same operation occurs when switching from recognition model M12 to M11.
[0075] Furthermore, if the switching becomes unnecessary while video is being input to both recognition models during the switching process (S121), the system may revert to the original recognition model. In other words, it is not necessary to switch to the target recognition model. If video is initially input to both recognition models in anticipation of a decrease in bitrate, but the situation changes and it is predicted that the bitrate will not change (or will recover immediately even if it decreases), the switching may be interrupted and the system may revert to the original recognition model. Note that the processing flow shown in Figure 11 is just one example, and the order of each process is not limited to this. Some processes may be executed in a different order, or some processes may be executed in parallel. For example, if terminal 100 or MEC400 is equipped with a prediction unit 230, a decision unit 240, a switching unit 250, and a storage unit 260, S116 to S118 may be executed between S111 and S112. Also, S116 to S118 may be executed in parallel with S111 to S115, provided that it is before input switching.
[0076] As described above, in this embodiment, the remote monitoring system predicts changes in the video bitrate and switches the recognition model that analyzes the video according to the predicted changes in bitrate. In addition, a frame from slightly before the switch is input to both the recognition model before and after the switch. This allows for the appropriate selection of the recognition model according to the changes in bitrate, and improves the recognition accuracy of the switched recognition model compared to simply switching the video input destination as shown in Figure 5.
[0077] (Embodiment 2) Next, Embodiment 2 will be described. In this embodiment, an example will be described in which video is input to the target recognition model using a buffer.
[0078] Figure 12 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 12, in this embodiment, the central server 200 is equipped with a buffer 270 in addition to the configuration of Embodiment 1. The other configurations are the same as in Embodiment 1. Here, we will mainly describe the configurations that differ from Embodiment 1.
[0079] Buffer 270 buffers the received video decoded by decoder 220. Buffer 270 holds the number of frames required by each recognition model for video recognition. It may hold the number of pre-input frames required for each recognition model, or it may hold the largest number of pre-input frames required by all recognition models.
[0080] When switching recognition models, the switching unit 250 retrieves frames held in the buffer 270 and inputs the received video, including the retrieved frames, to the new recognition model. The switching unit 250 retrieves the required number of pre-input frames from the buffer 270 for the new recognition model and inputs the received video, including the retrieved frames, to the new recognition model. For example, the buffer sizes of multiple buffers may be set to match the number of pre-input frames for each recognition model, and frames corresponding to the number of pre-input frames may be retrieved from the buffer corresponding to the recognition model. Alternatively, the video, including frames held in the buffer 270 at the switching timing, may be input to the new recognition model. In this case, it is not necessary to input the video from the pre-input timing as in Embodiment 1.
[0081] As described above, the remote monitoring system of Embodiment 1 may further include a buffer, and the frames held in the buffer may be input to the target recognition model. This makes it possible to improve the recognition accuracy of the target recognition model, similar to Embodiment 1.
[0082] (Embodiment 3) Next, Embodiment 3 will be described. In this embodiment, an example of switching the recognition model in accordance with changes in the video frame rate will be described.
[0083] Figure 13 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 13, in this embodiment, the center server 200 is equipped with a frame identification unit 280 in addition to the configuration of Embodiment 1. The other configurations are the same as in Embodiment 1. This embodiment may also be applied to Embodiment 2. Here, we will mainly describe the configurations that differ from Embodiment 1.
[0084] In this embodiment, recognition models M11 and M12 are recognition models that have learned video at different frame rates. Recognition model M11 has learned video at a first frame rate, and recognition model M12 has learned video at a second frame rate. For example, the first frame rate is a higher frame rate than the second frame rate, recognition model M11 is a model for high frame rates, and recognition model M12 is a model for low frame rates, but this is not limited to this. Note that the first frame rate and second frame rate are not limited to these, but may also be a first frame rate range and a second frame rate range.
[0085] Furthermore, the recognition model may learn and analyze video with a predetermined bitrate and frame rate. Multiple recognition models may learn and analyze video with different bitrate and frame rate combinations. In this case, the recognition model is selected and switched according to the bitrate and frame rate of the video.
[0086] The memory unit 260 stores a frame rate-recognition model table as an example of a video recognition environment-recognition model table, which associates frame rates with recognition models. Figure 14 shows a specific example of a frame rate-recognition model table. In this example, frame rate FR1 is associated with recognition model M11, and frame rate FR2 is associated with recognition model M12. Frame rates FR1 and FR2 correspond to the frame rates of the video learned by each recognition model. For example, frame rate FR1 is a high frame rate higher than frame rate FR2, and frame rate FR2 is a low frame rate lower than frame rate FR1.
[0087] The prediction unit 230 monitors the frame rate of the received video and predicts changes in the frame rate. For example, the prediction unit 230 obtains the frame rate included in the header of the encoded data. Rather than being limited to the header of the encoded data, the terminal 100 may send a packet containing the encoded data and frame rate to the center communication unit 210, and the prediction unit 230 may obtain the frame rate from the received packet. Based on the history of past frame rates obtained periodically, the prediction unit 230 extracts trends in frame rate transitions and predicts subsequent changes in the frame rate. Furthermore, if the terminal 100 is equipped with a prediction unit 230, it may predict changes in the frame rate based on instructions from the MEC 400 or the frame rate determined based on measurements by the communication quality measurement unit of the terminal 100.
[0088] The determination unit 240 determines the target recognition model and switching timing according to the predicted change in frame rate. The determination unit 240 refers to the frame rate-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted frame rate. In the example frame rate-recognition model table in Figure 14, if it is predicted that the frame rate will change from FR1 to FR2, it is decided to switch the recognition model from M11 to M12, and the timing at which the frame rate changes from FR1 to FR2 is determined as the switching timing. For example, the predicted frame rate is compared with FR1 and FR2, and the timing at which the predicted frame rate changes from a state close to FR1 to a state close to FR2 is set as the switching timing. If the frame rates FR1 and FR2 include a range of frame rates, it may be compared with the center of the range, or with any value in the range. Also, similar to Embodiment 1, the determination unit 240 determines the pre-input timing based on the number of pre-input frames corresponding to the target recognition model and the frame rate learned by the target recognition model.
[0089] The frame identification unit 280 identifies the frame interval, i.e., the frame rate, of the video to be input to the recognition model, according to the recognition model selected by the switching unit 250. The frame identification unit 280 identifies the frame interval, for example, by adjusting the frame interval. If the frame rate of the video input to the recognition model before and after switching is different, the frame identification unit 280 performs frame decimation or frame interpolation. Frame interpolation is the insertion of frames between video frames. The frame interval may be identified before the pre-input timing, from the pre-input timing to the switching timing, or after the switching timing. For example, the frame identification unit 280 refers to the frame rate-recognition model table in the storage unit 260, adjusts the frame interval of the input video based on the difference between the frame rate of the input video and the frame rate or frame rate range learned by the selected recognition model, and inputs the adjusted video to the recognition model. If the frame rate of the video is lower than the frame rate learned by the recognition model, frame interpolation is performed to match the frame rate learned by the recognition model. The method of frame interpolation is not limited. For example, the same frame may be inserted as the frame before or after the inserted frame, or a frame estimated based on changes in the image in past frames may be inserted. If the frame rate of the video is higher than the frame rate learned by the recognition model, frames are downsampled to match the frame rate learned by the recognition model. In addition, terminal 100 and MEC400 may be equipped with a frame identification unit 280, similar to the prediction unit 230, etc.
[0090] For example, suppose recognition model M11 is a recognition model that has learned video with a frame rate of 10fps, and recognition model M12 is a recognition model that has learned video with a frame rate of 30fps. In this case, when switching from recognition model M11 to M12 to input video with a frame rate of 10fps, the frame identification unit 280 performs frame interpolation on the input video and inputs the video with frame interpolation to 30fps to recognition model M12. Also, when switching from recognition model M12 to M11 to input video with a frame rate of 30fps, the frame identification unit 280 desamples frames from the input video and inputs the video with desampled frames to 10fps to recognition model M11.
[0091] As described above, in the remote monitoring system of Embodiment 1, the system may predict changes in the frame rate of the video and switch the recognition model that analyzes the video according to the predicted changes in the frame rate. This allows for the appropriate selection of the recognition model according to the changes in the frame rate, and, as in Embodiment 1, improves the recognition accuracy of the selected recognition model. Furthermore, by adjusting and specifying the input frame interval according to the recognition model, it is possible to input video with a frame rate suitable for the recognition model, thereby improving recognition accuracy.
[0092] (Embodiment 4) Next, Embodiment 4 will be described. In this embodiment, as a change in the video recognition environment, an example will be described in which the recognition model is switched in response to a change in the communication quality of the video being received.
[0093] Figure 15 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 15, in this embodiment, the center server 200 includes a communication quality measurement unit 290 in addition to the configuration of Embodiment 1. The other configurations are the same as in Embodiment 1. This embodiment may also be applied to other embodiments. For example, the recognition model M11 learns video of a first bitrate, and the recognition model M12 learns video of a second bitrate, similar to Embodiment 1. However, the recognition models M11 and M12 may learn video of different frame rates, similar to Embodiment 3. Furthermore, the recognition models M11 and M12 may learn video corresponding to different communication quality. Here, we will mainly describe the configurations that differ from Embodiment 1.
[0094] The communication quality measurement unit 290 measures the communication quality between the terminal 100 and the central server 200. Communication quality refers to the communication quality of the receiving path through which the central server 200 receives video from the terminal 100. Communication quality is, for example, communication speed, but may also be other indicators such as transmission delay or error rate. For example, the communication speed is measured based on the amount of data per unit time received by the central communication unit 210. Alternatively, the base station 300, terminal 100, or MEC 400 may be equipped with a communication quality measurement unit, and the communication quality measured or estimated by the communication quality measurement unit of the base station 300, terminal 100, or MEC 400 may be acquired.
[0095] The prediction unit 230 predicts changes in communication quality as a change in the image recognition environment. The prediction unit 230 periodically acquires the communication quality measured by the communication quality measurement unit 290, extracts the trend of communication quality transitions based on the acquired history of past communication quality, and predicts subsequent changes in communication quality. Figure 16 shows an example of predicting communication speed. As shown in Figure 16, future changes in communication speed are predicted from the history of past communication speed.
[0096] The determination unit 240 determines the target recognition model and switching timing in accordance with the predicted change in communication quality. If the recognition models M11 and M12 have learned video for each bitrate, the determination unit 240 determines the target recognition model and switching timing based on the bitrate corresponding to the communication quality. For example, the determination unit 240 estimates the bitrate of the received video from the predicted communication speed. Since the transmitting terminal 100 determines the bitrate according to the communication quality and encodes the video, the receiving center server 200 also determines the bitrate according to the communication quality in the same way as the terminal 100, thereby estimating the bitrate encoded by the terminal 100. For example, by associating the communication speed with the estimated bitrate, the bitrate can be estimated from the communication speed. The determination unit 240 determines the target recognition model and switching timing in accordance with the change in the estimated bitrate, similar to Embodiment 1. In the example in Figure 16, the switching timing is determined to be ts, when the bitrate changes to a predetermined value or less according to the communication speed. Also, similar to Embodiment 1, the pre-input timing ti is determined based on the switching timing. Furthermore, if recognition models M11 and M12 have learned video for each communication quality, the recognition model corresponding to the predicted communication quality will be used as the target recognition model.
[0097] As described above, in the remote monitoring system of Embodiment 1, changes in the communication quality for receiving video may be predicted, and the recognition model for analyzing the video may be switched according to the predicted changes in communication quality. This allows for the appropriate selection of the recognition model according to the changes in communication quality, and, as in Embodiment 1, improves the recognition accuracy of the switched-to recognition model.
[0098] (Embodiment 5) Next, Embodiment 5 will be described. In this embodiment, an example will be described in which the recognition model is switched according to the packet loss of the video packets as a communication quality included in the video recognition environment. The configuration of the remote monitoring system 1 according to this embodiment is the same as that shown in Figure 15 of Embodiment 4. Here, we will mainly describe the configuration that differs from Embodiment 4.
[0099] In this embodiment, recognition models M11 and M12 are recognition models that have learned video with different packet loss occurrences as examples of communication quality. For example, recognition model M11 has learned video without packet loss, and recognition model M12 has learned video with packet loss. Packet loss refers to the loss of all or some of the packets that transmit the data of a video frame because they cannot be properly received by the receiving side. This may be packet loss on a frame-by-frame basis, or packet loss over a predetermined period of time. Regardless of whether or not there is packet loss, recognition model M11 may learn video with a first packet loss rate, and recognition model M12 may learn video with a second packet loss rate. For example, the first packet loss rate may be lower than the second packet loss rate.
[0100] The memory unit 260 stores a packet loss-recognition model table as an example of a video recognition environment-recognition model table, which associates the occurrence of packet loss with the recognition model. Figure 17 shows a specific example of the packet loss-recognition model table. In this example, no packet loss is associated with recognition model M11, and packet loss is associated with recognition model M12. When associating the packet loss rate, a range of packet loss rates may also be associated.
[0101] The communication quality measurement unit 290 measures the occurrence of packet loss, i.e., whether or not packet loss occurs, as part of the communication quality. It monitors packets received by the central communication unit 210 and measures whether or not packets are missing in each frame.
[0102] The prediction unit 230 predicts the occurrence of packet loss. The prediction unit 230 periodically acquires the occurrence of packet loss measured by the communication quality measurement unit 290, extracts trends in packet loss based on the acquired past packet loss history, and predicts the occurrence of packet loss in the future.
[0103] The decision unit 240 determines the target recognition model and switching timing according to the predicted packet loss situation. The decision unit 240 refers to the packet loss-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted packet loss situation. In the example packet loss-recognition model table in Figure 17, if it is predicted that the situation will change from no packet loss to packet loss, it is decided to switch the recognition model from M11 to M12, and the timing at which the situation will change from no packet loss to packet loss is determined as the switching timing.
[0104] As described above, in the remote monitoring system of Embodiment 4, changes in the occurrence of packet loss in the packets receiving video may be predicted, and the recognition model for analyzing the video may be switched according to the predicted changes in the occurrence of packet loss. This allows for the appropriate selection of the recognition model according to the changes in the occurrence of packet loss, and, as in Embodiment 4, the recognition accuracy of the switched-to recognition model can be improved.
[0105] (Embodiment 6) Next, Embodiment 6 will be described. In this embodiment, as a change in the video recognition environment, an example will be described in which the recognition model is switched according to a change in the scene in which the video was captured.
[0106] Figure 18 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 18, in this embodiment, the center server 200 is equipped with a scene analysis unit 291 in addition to the configuration of Embodiment 1. The other configurations are the same as in Embodiment 1. Note that this embodiment may also be applied to other embodiments. Here, we will mainly describe the configurations that differ from Embodiment 1.
[0107] In this embodiment, recognition models M11 and M12 are recognition models that have learned from videos of different scenes. A scene refers to the progress of a construction project, the work being done by workers, and their work locations. For example, recognition model M11 has learned from videos of a first work process, and recognition model M12 has learned from videos of a second work process.
[0108] The memory unit 260 stores a scene-recognition model table, which associates scenes with recognition models, as an example of an image recognition environment-recognition model table. Figure 19 shows a specific example of a scene-recognition model table. In this example, work process A is associated with recognition model M11, and work process B is associated with recognition model M12.
[0109] The scene analysis unit 291 analyzes scenes in the video. For example, the scene analysis unit 291 analyzes scenes in the video based on the recognition results of the recognition model M11 or M12. When the recognition models M11 and M12 recognize work content from the video, the work content and work process may be associated in advance, and the work process may be determined from the recognized work content. The terminal 100 may also be equipped with a scene analysis unit 291. If the terminal 100 is equipped with a scene analysis unit 291, it may analyze the scenes in the video based on the video acquired by the video acquisition unit 110. For example, the terminal 100 may be equipped with an object detection unit, and the scene analysis unit 291 may analyze the scenes based on the correspondence information between objects detected by the object detection unit and the scenes.
[0110] The prediction unit 230 predicts changes in the video scene. The prediction unit 230 periodically acquires scenes analyzed by the scene analysis unit 291 and predicts subsequent scene changes based on the history of acquired past scenes. For example, it acquires schedule information for work processes and, based on the schedule information, predicts the completion of work, the content of the next work, and the next work process from the analyzed work content and work processes. The schedule information may include the time and content of each work process.
[0111] The decision unit 240 determines the target recognition model and switching timing according to the predicted scene change. The decision unit 240 refers to the scene-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted scene. In the example scene-recognition model table in Figure 19, if it is predicted that the work process will change from A to B, it is decided to switch the recognition model from M11 to M12, and the timing of the change from work process A to B is determined as the switching timing.
[0112] As described above, in the remote monitoring system of Embodiment 1, the system may predict changes in the scene in which the video has been captured and switch the recognition model that analyzes the video according to the predicted changes in the scene. This allows for the appropriate selection of the recognition model according to the changes in the scene, and, as in Embodiment 1, improves the recognition accuracy of the selected recognition model.
[0113] (Embodiment 7) Next, Embodiment 7 will be described. In this embodiment, as a change in the image recognition environment, an example will be described in which the recognition model is switched according to a change in the size of an object included in the image.
[0114] Figure 20 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 20, in this embodiment, the center server 200 is equipped with an object detection unit 292 in addition to the configuration of Embodiment 1. The other configurations are the same as in Embodiment 1. Note that this embodiment may also be applied to other embodiments. Here, we will mainly describe the configurations that differ from Embodiment 1.
[0115] In this embodiment, recognition models M11 and M12 are recognition models that have learned images of objects of different sizes. Recognition model M11 has learned images of a first object size, and recognition model M12 has learned images of a second object size. For example, the first object size is larger than the second object size, so recognition model M11 is a model for larger objects and recognition model M12 is a model for smaller objects, but this is not limited to this. The size of an object, or object dimensions, is the number of pixels in the area of the image in which the object is visible. For example, the closer an object is to the camera, the larger its size will be, and the further away an object is from the camera, the smaller its size will be. Also, the size of an object changes depending on the camera's zoom.
[0116] The memory unit 260 stores an object size-recognition model table as an example of an image recognition environment-recognition model table, which associates the size of an object with a recognition model. Figure 21 shows a specific example of an object size-recognition model table. In this example, size A is associated with recognition model M11, and size B is associated with recognition model M12. Sizes A and B may include a range of object sizes. Sizes A and B correspond to the object sizes of the images that each recognition model has learned; for example, size A is larger than size B, and size B is smaller than size A.
[0117] The object detection unit 292 detects objects in the video. For example, the object detection unit 292 extracts regions containing objects from each image of the video and detects objects within the extracted regions. The type of object to be recognized may be set in advance, and the size of the region of the object to be recognized may be extracted as the size of the object from among the detected objects. The object detection unit 292 may recognize objects in the image using an object recognition engine that uses machine learning. Alternatively, object detection results may be obtained from recognition models M11 or M12.
[0118] The prediction unit 230 predicts changes in the size of an object. The prediction unit 230 periodically acquires the size of the object detected by the object detection unit 292, extracts the trend of changes in the size of the object based on the acquired history of past object sizes, and predicts subsequent changes in the size of the object. For example, it tracks the target object between frames of the video, compares the sizes of the tracked objects, and predicts changes in size.
[0119] The determination unit 240 determines the target recognition model and the switching timing according to the predicted change in the size of the object. The determination unit 240 refers to the object size-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted size of the object. In the example object size-recognition model table in Figure 21, if it is predicted that the size of the object will change from size A to size B, it is decided to switch the recognition model from M11 to M12, and the timing at which the size changes from size A to size B is determined as the switching timing. For example, the predicted size of the object is compared with sizes A and B, and the timing at which the predicted size of the object changes from a state close to size A to a state close to size B is set as the switching timing. If sizes A and B include a range of sizes, they may be compared with the center of the range, or with any value in the range.
[0120] As described above, in the remote monitoring system of Embodiment 1, the system may predict changes in the size of objects included in the video and switch the recognition model that analyzes the video according to the predicted changes in the size of the objects. This allows for the appropriate selection of the recognition model according to changes in the size of objects, and, as in Embodiment 1, improves the recognition accuracy of the selected recognition model.
[0121] (Embodiment 8) Next, Embodiment 8 will be described. In this embodiment, as a change in the image recognition environment, an example will be described in which the recognition model is switched according to a change in the motion speed of an object included in the image.
[0122] Figure 22 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 22, in this embodiment, the center server 200 is equipped with a speed analysis unit 293 in addition to the configuration of Embodiment 1. The other configurations are the same as in Embodiment 1. Note that this embodiment may also be applied to other embodiments. Here, we will mainly describe the configurations that differ from Embodiment 1.
[0123] In this embodiment, recognition models M11 and M12 are recognition models that have learned images of objects moving at different speeds. Recognition model M11 has learned images of objects moving at a first speed, and recognition model M12 has learned images of objects moving at a second speed. The computational load of the recognition model also differs depending on the speed of the object to be recognized. For example, the first speed is lower than the second speed, so recognition model M11 is a low-computational-load model that can recognize only slow movements, and recognition model M12 is a high-computational-load model that can recognize even fast movements, but this is not limited to this. Note that it is not limited to the first speed and the second speed, but may also be a first speed range and a second speed range.
[0124] The memory unit 260 stores a motion speed-recognition model table as an example of an image recognition environment-recognition model table, which associates the motion speed of an object with a recognition model. Figure 23 shows a specific example of the motion speed-recognition model table. In this example, velocity A is associated with recognition model M11, and velocity B is associated with recognition model M12. Velocities A and B correspond to the motion speeds of the images learned by each recognition model; for example, velocity A is slower than velocity B, and velocity B is faster than velocity A.
[0125] The velocity analysis unit 293 analyzes the motion speed of objects in the video. For example, the velocity analysis unit 293 analyzes the motion speed based on the recognition results of the recognition model M11 or M12. When the recognition models M11 and M12 recognize an action, the action and motion speed may be associated in advance, and the motion speed may be determined from the recognized action. For example, if a person is walking or leveling the ground is recognized, it is determined to be a low-speed action, and if a person is running or throwing an object is recognized, it is determined to be a high-speed action. For example, an object in the video may be detected, the movement of the object between frames may be extracted, and the velocity may be determined from the extracted amount of movement. The terminal 100 may also be equipped with a speed analysis unit 293. If the terminal 100 is equipped with a speed analysis unit 293, the speed of motion in the video may be analyzed based on the video acquired by the video acquisition unit 110. For example, the terminal 100 may be equipped with an object detection unit, and the speed analysis unit 293 may analyze the motion speed based on the movement of an object detected by the object detection unit.
[0126] The prediction unit 230 predicts changes in the motion speed of an object. The prediction unit 230 periodically acquires the motion speed of the object analyzed by the velocity analysis unit 293, extracts the trend of the transition in the motion speed of the object based on the acquired history of past motion speeds, and predicts subsequent changes in the motion speed of the object.
[0127] The decision unit 240 determines the target recognition model and switching timing according to the predicted change in the object's operating speed. The decision unit 240 refers to the operating speed-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted operating speed of the object. In the example operating speed-recognition model table in Figure 23, if it is predicted that the object's operating speed will change from speed A to speed B, it is decided to switch the recognition model from M11 to M12, and the timing at which the speed changes from speed A to speed B is determined as the switching timing.
[0128] As described above, in the remote monitoring system of Embodiment 1, the system may predict changes in the motion speed of objects included in the video and switch the recognition model that analyzes the video according to the predicted changes in the motion speed of the objects. This allows for the appropriate selection of the recognition model according to changes in the motion speed of objects, enabling recognition of both slow and fast motion with the minimum necessary computational load, and, as in Embodiment 1, improving the recognition accuracy of the switched recognition model.
[0129] (Embodiment 9) Next, Embodiment 9 will be described. In this embodiment, as a change in the video recognition environment, an example of switching the recognition model in accordance with a change in the video shooting state will be described.
[0130] Figure 24 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 24, in this embodiment, the center server 200 is equipped with a state analysis unit 294 in addition to the configuration of Embodiment 1. The other configurations are the same as in Embodiment 1. Note that this embodiment may also be applied to other embodiments. Here, we will mainly describe the configurations that differ from Embodiment 1.
[0131] In this embodiment, recognition models M11 and M12 are models that have learned from videos shot under different conditions. These shooting conditions include fixed-camera shooting from a fixed position and mobile-camera shooting from a moving position. For example, recognition model M11 learns from videos shot using fixed-camera shooting, and recognition model M12 learns from videos shot using mobile-camera shooting. However, not limited to fixed-camera / mobile-camera shooting, recognition model M11 may learn from videos shot while moving at a first speed, such as slow movement, and recognition model M12 may learn from videos shot while moving at a second speed, such as high-speed movement.
[0132] The memory unit 260 stores a shooting state-recognition model table as an example of an image recognition environment-recognition model table, which associates the shooting state with the recognition model. Figure 25 shows a specific example of the shooting state-recognition model table. In this example, fixed shooting is associated with recognition model M11, and moving shooting is associated with recognition model M12. When associating the speed of movement, a range of speeds may also be associated.
[0133] The state analysis unit 294 analyzes the shooting state of the video. Based on the recognition results of the recognition model M11 or M12, the state analysis unit 294 may detect the shooting state, such as fixed shooting or moving shooting. For example, if the camera is an onboard vehicle camera and the video shows traffic lights at an intersection, the shooting state may be determined according to the color of the traffic lights in front. In the case of an onboard vehicle camera, the shooting state may also be detected according to vehicle control information or user operation information acquired from the vehicle. For example, the shooting state may be determined according to vehicle speed information, engine on / off status, and operation of the shift lever, brake pedal, and accelerator pedal. The terminal 100 may also include a state analysis unit 294. If the terminal 100 includes a state analysis unit 294, it may analyze the shooting state of the video based on the video acquired by the video acquisition unit 110. For example, the terminal 100 may include an object detection unit, and the state analysis unit 294 may analyze the shooting state based on the color and movement of the object detected by the object detection unit.
[0134] The prediction unit 230 predicts changes in the video shooting state. The prediction unit 230 periodically acquires the shooting state analyzed by the state analysis unit 294 and predicts subsequent changes in the shooting state based on the acquired history of past shooting states. For example, if fixed shooting / moving shooting is detected, the prediction unit 230 predicts changes between fixed shooting and moving shooting from the past history. Alternatively, if the color of the traffic light in front is detected, the prediction unit may estimate the vehicle's driving state by predicting that the traffic light color will change and predict changes between fixed shooting and moving shooting. If user operation information of the vehicle is detected, the prediction unit may estimate the vehicle's driving state by predicting the next user operation and predict changes between fixed shooting and moving shooting.
[0135] The determination unit 240 determines the target recognition model and switching timing according to the predicted change in the shooting state of the video. The determination unit 240 refers to the shooting state-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted shooting state. In the example shooting state-recognition model table in Figure 25, if it is predicted that the shooting state will change from fixed shooting to moving shooting, it is decided to switch the recognition model from M11 to M12, and the timing of the change from fixed shooting to moving shooting is determined as the switching timing. Alternatively, if the color of the traffic light in front is detected, the timing of the change from fixed shooting to moving shooting may be used as the timing of the change, and the target recognition model and switching timing may be determined accordingly. If the operation of the vehicle user is predicted, the timing of the start of accelerator pedal operation may be used as the timing of the change from fixed shooting to moving shooting, and the target recognition model and switching timing may be determined accordingly.
[0136] As described above, in the remote monitoring system of Embodiment 1, changes in the video shooting state, such as the start of camera movement, may be predicted, and the recognition model that analyzes the video may be switched according to the predicted changes in the shooting state. This allows for the appropriate selection of the recognition model according to changes in the video shooting state, and, as in Embodiment 1, the recognition accuracy of the switched-to recognition model can be improved.
[0137] (Embodiment 10) Next, Embodiment 10 will be described. In this embodiment, two recognition models are placed at different locations, and an example is described in which the recognition model is switched according to a change in the amount of computation of the video as a change in the video recognition environment.
[0138] Figure 26 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 26, the basic configuration of this embodiment is the same as that of Embodiment 1, but the arrangement of each part is different. Specifically, the MEC400 is equipped with the recognition model M11, and the center server 200 is equipped with the recognition model M12. In addition, the terminal 100 is equipped with a prediction unit 230, a decision unit 240, a switching unit 250, and a storage unit 260. Furthermore, the terminal 100 is equipped with a computation amount analysis unit 295. Note that this embodiment may also be applied to other embodiments. Here, we will mainly describe the configuration that differs from Embodiment 1.
[0139] In this embodiment, recognition models M11 and M12 are recognition models that have learned videos with different computing powers and different amounts of computation required for video analysis and recognition. Recognition model M11 learns videos that can be analyzed and recognized with a first amount of computation, and recognition model M12 learns videos that can be analyzed and recognized with a second amount of computation. For example, the first amount of computation is lower than the second amount of computation, so recognition model M11 is a low-computation-load model and recognition model M12 is a high-computation-load model, but it is not limited to this.
[0140] The memory unit 260 stores a computation amount-recognition model table as an example of an image recognition environment-recognition model table, which associates the computation amount of an analyzable and recognizable image with a recognition model. Figure 27 shows a specific example of a computation amount-recognition model table. In this example, computation amount A is associated with recognition model M11, and computation amount B is associated with recognition model M12. Computation amounts A and B may include a range of computation amounts. Computation amounts A and B correspond to the computation amounts of images learned by each recognition model; for example, computation amount A is a low computation amount lower than computation amount B, and computation amount B is a high computation amount higher than computation amount A.
[0141] The computational amount analysis unit 295 analyzes the computational amount necessary for video analysis and recognition. For example, the computational amount analysis unit 295 may associate objects with computational amounts, detect objects in the video, and determine the computational amount from the detected objects. Alternatively, it may detect objects in the video, extract the movement of objects between frames, and determine the computational amount from the extracted movement. Furthermore, it may associate actions recognized by recognition models M11 and M12 with computational amounts, obtain recognition results from recognition models M11 or M12, and determine the computational amount from the recognized actions.
[0142] The prediction unit 230 predicts changes in the amount of computation required for video analysis and recognition. The prediction unit 230 periodically acquires the amount of computation analyzed by the computation amount analysis unit 295 and predicts subsequent changes in the amount of computation based on the acquired history of past computation amounts.
[0143] The decision unit 240 determines the target recognition model and switching timing according to the predicted change in computation amount. The decision unit 240 refers to the computation amount-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted computation amount. In the example computation amount-recognition model table in Figure 27, if it is predicted that the computation amount will change from A to B, it is decided to switch the recognition model from M11 to M12, and the timing at which the computation amount changes from A to B is determined as the switching timing.
[0144] The switching unit 250 transmits video to the recognition model determined by the determination unit 240. If recognition model M11 is selected, the video is transmitted to the MEC400; if recognition model M12 is selected, the video is transmitted to the center server 200. The switching unit 250 switches the destination of the video transmission according to the switching timing. From the pre-input timing to the switching timing, the video is transmitted to both the recognition model before and after the switching timing; after the switching timing, the video is transmitted to the recognition model after the switching.
[0145] As described above, in the remote monitoring system of Embodiment 1, recognition models with different computational requirements may be placed at different locations. For example, by running the low-computational-requirement model on the MEC and the high-computational-requirement model on the center, the computing resources of both the MEC and the center can be efficiently utilized, increasing the total number of images that can be analyzed and recognized by the entire system. Furthermore, the recognition results obtained by the MEC's recognition model may be used on the terminal side or in the field. Since the MEC is often closer to the field than the central office, the MEC can transmit the recognition results to the terminal or field equipment faster. As a result, in this embodiment, by utilizing the MEC's recognition model, the recognition results can be used quickly on the terminal side or in the field.
[0146] (Embodiment 11) Next, Embodiment 11 will be described. In this embodiment, two recognition models are placed at different locations, and an example is described in which the recognition model is switched according to a change in the bandwidth used to transmit the video, as a change in the video recognition environment.
[0147] Figure 28 shows an example configuration of the remote monitoring system 1 according to this embodiment. As shown in Figure 28, in this embodiment, compared to embodiment 10, the terminal 100 is equipped with a bandwidth acquisition unit 296 instead of a computation amount analysis unit 295. The other configurations are the same as in embodiment 10. Here, we will mainly describe the configurations that differ from embodiment 10. In this embodiment, the recognition models M11 and M12 may be recognition models with different computation amounts, as in embodiment 10, or they may be the same recognition model.
[0148] The memory unit 260 stores a transmission bandwidth-recognition model table as an example of a video recognition environment-recognition model table, which associates the transmission bandwidth between the terminal and the central server with the recognition model. Figure 29 shows a specific example of a transmission bandwidth-recognition model table. In this example, transmission bandwidth A is associated with recognition model M11, and transmission bandwidth B is associated with recognition model M12. Transmission bandwidth A and transmission bandwidth B have different bandwidths. For example, transmission bandwidth A is a narrow bandwidth, which is narrower than transmission bandwidth B, and transmission bandwidth B is a wide bandwidth, which is wider than transmission bandwidth A.
[0149] The bandwidth acquisition unit 296 acquires the transmission bandwidth between the terminal 100 and the central server 200. The transmission bandwidth may also be determined based on the communication speed estimated based on the amount of data transmitted from the terminal communication unit 130. Alternatively, the communication speed measured by the base station 300 or the terminal 100 may be acquired, and the transmission bandwidth may be determined from the acquired communication speed.
[0150] The prediction unit 230 predicts changes in the transmission bandwidth. The prediction unit 230 periodically acquires the transmission bandwidth acquired by the bandwidth acquisition unit 296, extracts the trend of transmission bandwidth transitions based on the acquired history of past transmission bandwidth, and predicts subsequent changes in the transmission bandwidth.
[0151] The decision unit 240 determines the target recognition model and switching timing according to the predicted change in transmission bandwidth. The decision unit 240 refers to the transmission bandwidth-recognition model table in the storage unit 260 and determines the recognition model corresponding to the predicted transmission bandwidth. In the example transmission bandwidth-recognition model table in Figure 29, if it is predicted that the transmission bandwidth will change from A to B, it is decided to switch the recognition model from M11 to M12, and the timing of the change from transmission bandwidth A to transmission bandwidth B is determined as the switching timing.
[0152] As described above, in the remote monitoring system of Embodiment 10, two recognition models may be placed at different locations, and the recognition models may be switched according to changes in the transmission bandwidth. If the network bandwidth between the site and the center is sufficient, video recognition of the recognition model may be performed at the center; otherwise, video recognition of the recognition model may be performed at the MEC. This prevents a decrease in analysis accuracy due to the analysis of low-resolution video at the center. In addition, higher quality video can be transmitted to the recognition model on the MEC or the center side, improving recognition accuracy compared to when the recognition model is located in only one place.
[0153] This disclosure is not limited to the embodiments described above, and may be modified as appropriate without departing from its spirit.
[0154] Each configuration in the above-described embodiment may consist of hardware, software, or both, and may consist of one piece of hardware or software, or multiple pieces of hardware or software. Each device and each function (process) may be realized by a computer 30 having a processor 31 such as a CPU (Central Processing Unit) and a memory 32 as a storage device, as shown in Figure 30. For example, a program for performing the method (image processing method) in the embodiment may be stored in the memory 32, and each function may be realized by executing the program stored in the memory 32 with the processor 31.
[0155] These programs, when loaded into a computer, include a set of instructions (or software code) for causing the computer to perform one or more of the functions described in the embodiments. The programs may be stored on non-temporary computer-readable media or tangible storage media. Examples, but not limited to, include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drives (SSDs), or other memory technologies, CD-ROMs, digital versatile discs (DVDs), Blu-ray® discs, or other optical disc storage, magnetic cassettes, magnetic tapes, magnetic disk storage, or other magnetic storage devices. The programs may be transmitted over temporary computer-readable media or communication media. Examples, but not limited to, include electrical, optical, acoustic, or other forms of propagating signals.
[0156] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be understood by those skilled in the art within the scope of the present disclosure.
[0157] Some or all of the above embodiments may also be described as follows, but are not limited to the following: (Note 1) A first video analysis model that analyzes video corresponding to a first video recognition environment, A second video analysis model that analyzes video corresponding to a second video recognition environment, The system includes a switching means that switches the video analysis model used to analyze the video input data from the first video analysis model to the second video analysis model in response to a change in the video input data from the first video recognition environment to the second video recognition environment, The switching means inputs video input data, including data from a predetermined period prior to the switching timing, to the second video analysis model in response to a change from the first video recognition environment to the second video recognition environment in the video input data. Video processing system. (Note 2) The video input data, which includes data prior to the aforementioned switching timing, is video input data that includes frame number data used by the second video analysis model for video recognition. The video processing system described in Appendix 1. (Note 3) The switching means inputs the video input data of the frame count to both the first and second video analysis models. The video processing system described in Appendix 2. (Note 4) The system includes a prediction means for predicting changes in the video recognition environment in the aforementioned video input data, The switching means switches the video analysis model in accordance with the predicted changes in the video recognition environment. A video processing system as described in any one of the items 1 to 3 of the appendix. (Note 5) The aforementioned video recognition environment includes video parameters that indicate the quality of the video, A video processing system as described in any one of the items 1 to 4 of the appendix. (Note 6) The aforementioned video parameters include the frame rate, The system includes a means for identifying the frame interval of the video input data according to a video analysis model that inputs the video input data. The video processing system described in Appendix 5. (Note 7) The system includes a receiving means for receiving the aforementioned video input data via a network, The aforementioned video recognition environment includes the communication quality of the video input data received by the receiving means, A video processing system as described in any one of the items 1 to 6 of the appendix. (Note 8) The aforementioned video recognition environment includes the scene in which the video was captured, the size of the objects included in the video, the speed of movement of the objects included in the video, or the shooting conditions under which the video was captured. A video processing system as described in any one of the items 1 to 7 of the appendix. (Note 9) The first video analysis model is located at either the edge or the cloud. The second video analysis model is located at the edge and the other of the cloud, A video processing system as described in any one of the items 1 to 8 of the appendix. (Note 10) A first video analysis model that analyzes video corresponding to a first video recognition environment, A second video analysis model that analyzes video corresponding to a second video recognition environment, The system includes a switching means that switches the video analysis model used to analyze the video input data from the first video analysis model to the second video analysis model in response to a change in the video input data from the first video recognition environment to the second video recognition environment, The switching means inputs video input data, including data from a predetermined period prior to the switching timing, to the second video analysis model in response to a change from the first video recognition environment to the second video recognition environment in the video input data. Image processing device. (Note 11) The video input data, which includes data from a predetermined period prior to the aforementioned switching timing, is video input data that includes frame number data used by the second video analysis model for video recognition. The video processing device described in Appendix 10. (Note 12) The switching means inputs the video input data of the frame count to both the first and second video analysis models. The video processing device described in Appendix 11. (Note 13) The system includes a prediction means for predicting changes in the video recognition environment in the aforementioned video input data, The switching means switches the video analysis model in accordance with the predicted changes in the video recognition environment. An image processing device as described in any one of the appendices 10 to 12. (Note 14) The aforementioned video recognition environment includes video parameters that indicate the quality of the video, The image processing device described in any one of the appendices 10 to 13. (Note 15) The aforementioned video parameters include the frame rate, The system includes a means for identifying the frame interval of the video input data according to a video analysis model that inputs the video input data. The video processing device described in Appendix 14. (Note 16) In response to the change from a first video recognition environment to a second video recognition environment in the input video data, the video analysis model that analyzes the video input data is switched from a first video analysis model that analyzes video corresponding to the first video recognition environment to a second video analysis model that analyzes video corresponding to the second video recognition environment. In response to the change from the first video recognition environment to the second video recognition environment in the aforementioned video input data, video input data including data from a predetermined period prior to the switching timing is input to the second video analysis model. Image processing methods. (Note 17) The video input data, which includes data from a predetermined period prior to the aforementioned switching timing, is video input data that includes frame number data used by the second video analysis model for video recognition. The video processing method described in Appendix 16. (Note 18) The video input data of the aforementioned number of frames is input to both the first and second video analysis models. The video processing method described in Appendix 17. (Note 19) Predicting changes in the video recognition environment in the aforementioned video input data, In response to the predicted changes in the image recognition environment, the image analysis model is switched. The video processing method described in any one of the appendices 16 to 18. (Note 20) The aforementioned video recognition environment includes video parameters that indicate the quality of the video, The video processing method described in any one of the appendices 16 to 19. (Note 21) The aforementioned video parameters include the frame rate, Depending on the video analysis model that inputs the aforementioned video input data, the frame interval of the video input data is identified. The video processing method described in Appendix 20. (Note 22) In response to the change from a first video recognition environment to a second video recognition environment in the input video data, the video analysis model that analyzes the video input data is switched from a first video analysis model that analyzes video corresponding to the first video recognition environment to a second video analysis model that analyzes video corresponding to the second video recognition environment. In response to the change from the first video recognition environment to the second video recognition environment in the aforementioned video input data, video input data including data from a predetermined period prior to the switching timing is input to the second video analysis model. A video processing program that instructs a computer to perform a particular task. [Explanation of Symbols]
[0158] 1. Remote monitoring system 10. Video Processing System 11 Switching section 20, 21, 22 Video Processing Devices 30 Computers 31 processors 32 memory 100 devices 101 Camera 102 Compression Efficiency Optimization Function 110 Video Acquisition Unit 120 encoders 130 Terminal Communication Unit 200 Center Servers 201 Video Recognition Function 202 Alert generation function 203 GUI drawing function 204 Screen display function 210 Center Communications Department 220 Decoders 230 Prediction Section 240 Decision Section 250 Switching section 260 Storage section 270 buffers 280 Frame Identification Section 290 Communication Quality Measurement Unit 291 Scene Analysis Department 292 Object detection unit 293 Speed analysis section 294 Condition Analysis Department 295 Computation amount analysis section 296 Bandwidth Acquisition Section 300 base stations 400 MEC 401 Compression Bitrate Control Function M1, M2, M11, M12 Recognition Models
Claims
1. A first video analysis model that analyzes video corresponding to a first video recognition environment, A second video analysis model that analyzes video corresponding to a second video recognition environment, The system includes a switching means that switches the video analysis model used to analyze the video input data from the first video analysis model to the second video analysis model in response to a change in the video input data from the first video recognition environment to the second video recognition environment, The switching means inputs video input data, including data from a predetermined period prior to the switching timing, to the second video analysis model in response to a change from the first video recognition environment to the second video recognition environment in the video input data. Video processing system.
2. The video input data, which includes data from a predetermined period prior to the aforementioned switching timing, is video input data that includes frame number data used by the second video analysis model for video recognition. The image processing system according to claim 1.
3. The switching means inputs the video input data of the frame count to both the first and second video analysis models. The image processing system according to claim 2.
4. The system includes a prediction means for predicting changes in the video recognition environment in the aforementioned video input data, The switching means switches the video analysis model in accordance with the predicted changes in the video recognition environment. The video processing system according to any one of claims 1 to 3.
5. The aforementioned video recognition environment includes video parameters that indicate the quality of the video, The video processing system according to any one of claims 1 to 3.
6. The aforementioned video parameters include the frame rate, The system includes a means for identifying the frame interval of the video input data according to a video analysis model that inputs the video input data. The image processing system according to claim 5.
7. The system includes a receiving means for receiving the aforementioned video input data via a network, The aforementioned video recognition environment includes the communication quality of the video input data received by the receiving means, The video processing system according to any one of claims 1 to 3.
8. The aforementioned video recognition environment includes the scene in which the video was captured, the size of the objects included in the video, the speed of movement of the objects included in the video, or the shooting conditions under which the video was captured. The video processing system according to any one of claims 1 to 3.
9. A first video analysis model that analyzes video corresponding to a first video recognition environment, A second video analysis model that analyzes video corresponding to a second video recognition environment, The system includes a switching means that switches the video analysis model used to analyze the video input data from the first video analysis model to the second video analysis model in response to a change in the video input data from the first video recognition environment to the second video recognition environment, The switching means inputs video input data, including data from a predetermined period prior to the switching timing, to the second video analysis model in response to a change from the first video recognition environment to the second video recognition environment in the video input data. Image processing device.
10. In response to the change from a first video recognition environment to a second video recognition environment in the input video data, the video analysis model that analyzes the video input data is switched from a first video analysis model that analyzes video corresponding to the first video recognition environment to a second video analysis model that analyzes video corresponding to the second video recognition environment. In response to the change from the first video recognition environment to the second video recognition environment in the aforementioned video input data, video input data including data from a predetermined period prior to the switching timing is input to the second video analysis model. Image processing methods.