Real-time video analytics methods and systems based on dynamic detection and cloud-based AI

CN119629388BActive Publication Date: 2026-08-14CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-02
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本发明提供一种基于动态检测与云端AI结合的实时视频分析方法和系统,以解决现有技术中因仅利用摄像头端侧的单一AI算法进行数据分析而极大地限制了同时处理能力,且无法满足日益复杂多变的监控与分析需求,现有方法因从摄像头检测到事件到视频数据完整传输至云端才开始分析而导致分析时间差,进而导致分析响应存在延迟,严重影响了实时监控效率,甚至还会导致错过关键事件的即时处理的问题,资源成本高等的技术问题,本发明要解决的技术问题通过以下技术方案来实现

Benefits of technology

[0011]与现有技术相比,本发明基于监控摄像头动态检测诸如指定区域面积变化、光线亮度变化等多种画面变化并截图传输给云端AI实时分析,根据所构建的背景模型进行画面变化判断以确定是否触发联动设备以进行相应联动检测,并向所述当前监控摄像头发送所述画面截图所对应的返回信息,进一步将所述画面截图所对应的返回信息、当前监控摄像头的信息、画面变化的时间上报到调度服务器,以调度到一个或多个相应分析服务器进行分析,并将分析结果通知视频APP,有效解决了摄像头对光线亮度微小变化感知慢检测延迟的问题。此外,将摄像头与FTTR设备联动,有效解决了摄像头放置角度检测不到人、动物或其他物体的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119629388B_ABST
    Figure CN119629388B_ABST
Patent Text Reader

Abstract

This invention relates to the field of video processing technology. It provides a real-time video analysis method and system based on dynamic detection and cloud-based AI. The method includes: uploading real-time captured video clips from a current surveillance camera to object storage; taking screenshots based on the captured video clips; judging image changes based on a constructed background model, specifically including changes in area and light intensity; determining whether to trigger a linkage device for corresponding linkage detection; sending the return information corresponding to the screenshot to the current surveillance camera; reporting the return information corresponding to the screenshot, the information of the current surveillance camera, and the time of the image change to a scheduling server for scheduling to one or more corresponding analysis servers for analysis; and notifying the video app of the analysis results. This invention effectively solves the problem of slow detection delay in cameras sensing minute changes in light intensity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of video processing technology, and provides a real-time video analysis method and system based on the combination of dynamic detection and cloud AI. Background Technology

[0002] In existing video surveillance systems, AI analysis of video content typically employs two methods: real-time stream frame extraction analysis, which, while highly real-time, consumes significant computational resources and is extremely costly; and recording video first, then extracting and analyzing footage from the beginning and end of an event after it occurs, which reduces computational costs but suffers from poor real-time performance. Therefore, a video analysis method that can guarantee both real-time performance and effectively control analysis costs is needed.

[0003] Existing technologies suffer from the following problems: Limited AI algorithm support: Due to the limited computing power of cameras, traditional solutions often rely on a single AI algorithm at the camera's edge for data analysis. This severely restricts the ability to handle multiple AI functions simultaneously, failing to meet increasingly complex and dynamic monitoring and analysis needs. High resource costs: Real-time video streaming analysis requires the continuous processing of massive amounts of real-time data, placing extremely high demands on servers, storage devices, and network bandwidth. Furthermore, the procurement, deployment, and maintenance costs of high-performance hardware are high, increasing the overall operational burden and resulting in significant analysis latency. Current technologies rely on capturing video clips from cameras and uploading them to the cloud for AI analysis. This process, from camera detection of an event to the complete transmission of video data to the cloud and the commencement of analysis, often involves a time lag, leading to a significant delay in analysis response. This delay not only affects the efficiency of real-time monitoring but may also cause missed opportunities for timely handling of critical events.

[0004] Therefore, it is necessary to provide a real-time video analysis method based on the combination of dynamic detection and cloud AI to solve the above problems. Summary of the Invention

[0005] This invention provides a real-time video analysis method and system based on dynamic detection and cloud AI, which solves the problems of existing technologies that greatly limit simultaneous processing capabilities due to the use of a single AI algorithm on the camera side for data analysis, and cannot meet the increasingly complex and ever-changing monitoring and analysis needs. Existing methods cause analysis time lag because analysis only begins after the camera detects an event and the video data is completely transmitted to the cloud, resulting in delayed analysis response, which seriously affects real-time monitoring efficiency and may even lead to missing the immediate processing of critical events. The invention also addresses the high resource costs. The technical problems to be solved by this invention are achieved through the following technical solutions.

[0006] The first aspect of this invention proposes a real-time video analysis method based on dynamic detection and cloud AI. The real-time video analysis method includes: receiving a connection address request from a current surveillance camera; after authorization and authentication, providing the current surveillance camera with an available connection address to establish a long-term connection between the current surveillance camera and a streaming media server; uploading real-time captured video segments to object storage; taking screenshots based on the captured video segments; and judging image changes based on a constructed background model, specifically including changes in area and light intensity; determining whether to trigger a linkage device for corresponding linkage detection based on the image change judgment result, and sending the return information corresponding to the screenshot to the current surveillance camera; reporting the return information corresponding to the screenshot, the information of the current surveillance camera, and the time of the image change to a scheduling server for scheduling to one or more corresponding analysis servers for analysis; receiving the analysis results from one or more corresponding analysis servers and notifying the video app of the analysis results, wherein the analysis results include AI recognition analysis results, AI events, and event start times.

[0007] The second aspect of this invention proposes a real-time video analysis system based on the combination of dynamic detection and cloud AI, which is consistent with the real-time video analysis method based on the combination of dynamic detection and cloud AI described in the first aspect of this invention. The real-time video analysis system includes: a receiving and processing module, used to receive a connection address request from a current surveillance camera; and, upon successful authorization and authentication, to provide an available connection address to the current surveillance camera, enabling the current surveillance camera to establish a long-term connection with a streaming media server; and a scene change judgment module, used to upload real-time captured video segments to object storage, perform scene screenshots based on the real-time captured video segments, and draw based on a constructed background model. The system includes a scene change judgment module, which specifically includes changes in area and light intensity; a determination module, which determines whether to trigger the linkage device to perform corresponding linkage detection based on the scene change judgment result, and sends the return information corresponding to the scene screenshot to the current monitoring camera; a scheduling module, which reports the return information corresponding to the scene screenshot, the information of the current monitoring camera, and the time of the scene change to the scheduling server, so as to schedule it to one or more corresponding analysis servers for analysis; and a notification module, which receives the analysis results of one or more corresponding analysis servers and notifies the video APP of the analysis results, including AI recognition analysis results, AI events, and event start time.

[0008] In a third aspect of the present invention, an electronic device is provided, including: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the real-time video analysis method based on the combination of dynamic detection and cloud AI described in the first aspect of the present invention.

[0009] In a fourth aspect of the present invention, a computer-readable medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the real-time video analysis method based on the combination of dynamic detection and cloud AI described in the first aspect of the present invention is implemented.

[0010] The embodiments of the present invention include the following advantages:

[0011] Compared with the prior art, the present invention dynamically detects various picture changes such as the change of the specified area area and the change of the light brightness by the monitoring camera and captures and transmits the screenshots to the cloud AI for real-time analysis, determines whether to trigger the linkage device for corresponding linkage detection according to the constructed background model, and sends the return information corresponding to the picture screenshot to the current monitoring camera, and further reports the return information corresponding to the picture screenshot, the information of the current monitoring camera, and the time of the picture change to the scheduling server to schedule to one or more corresponding analysis servers for analysis, and notifies the video APP of the analysis result, effectively solving the problem that the camera has a slow perception of small changes in light brightness and a detection delay. In addition, by linking the camera with the FTTR device, the problem that the camera placement angle cannot detect people, animals or other objects is effectively solved.

[0012] In addition, by deploying the AI algorithm analysis server and the object storage in the same intranet, the problem of a large amount of downloaded pictures occupying bandwidth is effectively solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a flowchart of the steps of the real-time video analysis method based on the combination of dynamic detection and cloud AI of the present invention;

[0014] Figure 2 is a partial flowchart of a specific embodiment of the real-time video analysis method based on the combination of dynamic detection and cloud AI of the present invention;

[0015] Figure 3 is another partial process of a specific embodiment of the real-time video analysis method based on the combination of dynamic detection and cloud AI of the present invention;

[0016] Figure 4 is a structural block diagram of the real-time video analysis system based on the combination of dynamic detection and cloud AI of the present invention;

[0017] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention;

[0018] Figure 6 This is a schematic diagram of a computer-readable medium embodiment according to the present invention. Detailed Implementation

[0019] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] In view of the above problems, this invention proposes a real-time video analysis method based on dynamic detection and cloud AI. This method relies on a surveillance camera to dynamically detect various scene changes, such as changes in the area of ​​a specified region and changes in light intensity, and transmits screenshots to cloud AI for real-time analysis. Based on a constructed background model, it judges the scene changes to determine whether to trigger a linkage device for corresponding linkage detection, and sends the return information corresponding to the screenshot to the current surveillance camera. Furthermore, the return information corresponding to the screenshot, the current surveillance camera information, and the time of the scene change are reported to a scheduling server for scheduling to one or more corresponding analysis servers for analysis. The analysis results are then notified to the video app, effectively solving the problem of slow detection delay in the camera's perception of minute changes in light intensity. In addition, linking the camera with an FTTR device effectively solves the problem of the camera's placement angle failing to detect people, animals, or other objects.

[0021] It should be noted that the method of the present invention has a wide range of applications, and is particularly suitable for home indoor monitoring scenarios, indoor monitoring scenarios in large venues, etc.

[0022] Example 1

[0023] The following reference Figure 1 , Figure 2 , Figure 3 The present invention will be described in detail below.

[0024] Figure 1 This is a flowchart of the steps of the real-time video analysis method based on the combination of dynamic detection and cloud AI of the present invention. Figure 2 This is a partial flowchart of a specific implementation of the real-time video analysis method based on the combination of dynamic detection and cloud AI of the present invention. Figure 3 This is a partial flowchart illustrating another specific embodiment of the real-time video analysis method based on the combination of dynamic detection and cloud AI of the present invention.

[0025] exist Figure 2Application examples include surveillance cameras, streaming media servers, object storage, AI algorithm scheduling servers, AI algorithm analysis servers, and video apps. Surveillance cameras, for example, are used in home monitoring. They capture facial features, moving objects, object actions, and unusual scenes (such as identifying family members and visitors, pet movement in the home, falls or prolonged periods of inactivity by the elderly, and environmental changes such as smoke or potential hazards) from video images and transmit this data to object storage or streaming media servers to achieve remote monitoring, video playback, and alarm triggering. Object storage is a technology that stores data in the form of objects. In object storage, each object contains data and its metadata, and has a unique identifier for retrieving objects without knowing the physical location of the data. In this example, it is mainly used to store video and screenshot data captured by surveillance cameras.

[0026] Specifically, the streaming media server assigns connectable IP addresses to the surveillance cameras. After connecting to the network via methods such as Ethernet cable or Wi-Fi, the cameras establish a persistent communication connection with the streaming media server. The AI ​​algorithm scheduling server distributes analysis to different AI algorithm analysis servers based on predefined AI algorithm configurations or the user's subscribed AI package. The video app is a tool provided to users, allowing them to view live and replayed footage from the surveillance cameras and receive alarm notifications generated by the AI ​​analysis server.

[0027] It should be noted that in this invention, the AI ​​algorithm analysis server is equipped with a high-performance GPU and various AI algorithm models, such as Convolutional Neural Networks (CNN) for face recognition and object detection (e.g., pets, plants, electric vehicles), Recurrent Neural Networks (RNN) for human behavior analysis (e.g., falls, walking), and Random Forest algorithms for environmental monitoring (e.g., smoke). This enables efficient performance of specific analysis tasks such as smoke detection, mask-wearing detection, and electric vehicle recognition. The AI ​​algorithm analysis server and the object storage system are deployed on the same intranet, i.e., within the same intranet environment. Devices within this intranet typically have private IP addresses and are interconnected via network devices such as routers and switches, ensuring data transmission occurs only within the intranet without requiring an external network. This network connects the server and storage devices through routers, switches, and other network devices. Each server has a private IP address, and servers directly access stored screenshot data via their private IP addresses provided by a high-speed local area network (LAN), effectively reducing bandwidth consumption and improving analysis efficiency.

[0028] Below, in conjunction with Figure 2 The method of this invention will be further described in detail below, along with specific analytical tasks. The main flow of this method is as follows:

[0029] In step S101, a request for a connection address from the current surveillance camera is received. After authorization and authentication, an available connection address is provided to the current surveillance camera so that the current surveillance camera and the streaming media server can establish a long connection.

[0030] In one specific implementation, for the specific analysis task of electric vehicle identification, the current surveillance camera a connects to the network via Ethernet cable, Wi-Fi, or other means, and sends a request for a connection IP address to the streaming media server.

[0031] When the streaming media server receives the application request, it first performs authorization authentication.

[0032] In one specific implementation, token-based authorization is used. Specifically, the user first logs in using the video app with a username or password, and the server verifies the login. If the verification is successful, a token (e.g., represented by a token) is generated and returned to the video app.

[0033] The video app's local storage token. Upon first use, the app associates the camera's serial number with the camera owner (i.e., the user). The camera uses this token for verification in subsequent communications. For example, it's used to authorize application requests and determine if authorization has been successful.

[0034] In another implementation, when the current surveillance camera a accesses the network, an authorization authentication process is initiated.

[0035] Specifically, the current surveillance camera 'a' connects to the network via Ethernet cable, Wi-Fi, etc., sends a connection request to the streaming media server, and appends a token to the Authorization field of the HTTP header, for example, in the format of Bearer. <token>It also sends the serial number of the current surveillance camera a as a request parameter.

[0036] When the streaming media server receives a request, it extracts the token from the HTTP header and the sequence number from the request parameters, and verifies the validity of the token, including the signature, issuer, expiration time, etc.

[0037] Specifically, the streaming media server queries the database based on the extracted serial number to verify the legitimacy of the current surveillance camera 'a', including whether it is registered and authorized by the user, in order to complete the authorization and authentication process.

[0038] Furthermore, upon successful authorization or authentication, or after successful authorization, a usable connection address (e.g., an IP address) is provided to the current surveillance camera a, enabling the current surveillance camera a to establish a persistent connection with the streaming media server. Conversely, if authorization fails, the streaming media server rejects the connection request from the current surveillance camera a and returns a corresponding error message to the current surveillance camera a.

[0039] For example, the current surveillance camera 'a' uses an available connection address (e.g., hio.***.com / ?Src=¥……&%&*=3) to establish a long connection with the streaming media server for data transmission.

[0040] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.

[0041] Next, in step S102, the real-time captured video clips are uploaded to object storage, screenshots are taken based on the real-time captured video clips, and the changes in the scene are judged according to the constructed background model, specifically including changes in area and changes in light brightness.

[0042] Specifically, the video clips captured in real time by the surveillance camera are uploaded to object storage. Screenshots are also taken from the captured video clips.

[0043] In one specific implementation, for example, the current camera uploads a screenshot (e.g., a screenshot of the current page) every 5 seconds.

[0044] A background model is constructed, which is used to determine image changes. Specifically, a set of background grayscale images corresponding to each color image captured by the surveillance camera under unobstructed conditions is constructed, and the set includes multiple background grayscale images.

[0045] It's important to note that the background model includes a static image of the monitored scene when there is no movement—a grayscale image of the background. The surveillance camera captures the current monitored scene in real time and compares it with the background grayscale image. If there is a significant difference between the current scene and the background grayscale image, these differences are identified as changes in the monitored scene. If the changes in the monitored scene reach a preset threshold (such as area or speed), the camera considers the monitored scene to have changed. For example, someone walking by, the camera being moved, or a significant change in the brightness of the current scene will all be detected as changes in the monitored scene. After a change in the monitored scene, the camera will take a screenshot of the current page and upload the screenshot to object storage.

[0046] The process of judging changes in the image specifically includes the following steps:

[0047] Step S201: Acquire multiple frames of color images from the surveillance camera under unobstructed conditions, convert the multiple frames of color images into multiple background grayscale images, and calculate the average grayscale value of the same pixel position in the multiple background grayscale images using the following formula:

[0048]

[0049] Where mean(g)xy represents the average gray value of all pixel positions in multiple background grayscale images; It is the gray value of the background grayscale image corresponding to the i-th color image at pixel position (x, y), where x represents the horizontal coordinate of pixel position (x, y), y represents the vertical coordinate of pixel position (x, y), and i is a positive integer, specifically 1, 2, ..., n, where n represents the number of background grayscale images.

[0050] For example, after a surveillance camera is installed and started, multiple frames of color images from the camera under unobstructed conditions are acquired, and these color images are converted into corresponding background grayscale images, denoted as g = {g i |i∈[1,n]}, where g i This represents the i-th background grayscale image in the set, where i is the index of the grayscale image in the set, and n represents the number of background grayscale images.

[0051] Each background grayscale image g i Each background grayscale image contains multiple pixels, each with a grayscale value. For example, if the image resolution is W*H (width multiplied by height), then each background grayscale image has W*H pixels, and each pixel location has a grayscale value. For the same location (x, y) in each background grayscale image, the grayscale values ​​of all images gi are added together, and then divided by the total number of background grayscale images n to obtain the average grayscale value at that pixel location.

[0052] Step S202: Cut the real-time captured current image and the background grayscale image into regions of the same size in the horizontal and vertical directions. For each region, calculate the mean and standard deviation of the grayscale of the current image, and calculate the mean and standard deviation of the grayscale of the background grayscale image.

[0053] Specifically, after the surveillance camera is operating normally, multiple frames of color images are acquired at equal time intervals and converted into a background grayscale image, for example, denoted as g. ‘ ={g i |i∈[1,k]}

[0054] Combine the background grayscale image m and the real-time image g ‘ Divide the surface into regions of equal size in both the horizontal and vertical directions, denoted as {m}. ij |i∈[1, p], j∈[1, q]} and {a ij |i∈[1, p], j∈[1, q]}. For each region a ij Calculation region a ij mean gray level (a) ij ) and standard deviation std(a ij mean(a) ij ) is the region a ij The mean and standard deviation of all grayscale values ​​represent the distribution range of these grayscale values. For each region m... ij , computational region m ij mean gray level (m) ij ) and standard deviation std(m ij mean(m) ij ) is the region m ij The mean and standard deviation of all grayscale values ​​represent the range of these grayscale values.

[0055] Step S203: Based on the segmented regions, determine whether the current image has undergone a change in appearance. Specifically, determine that the current image has undergone a change in appearance when the proportion of regions in the current image that have undergone a change in appearance to the total number of regions is greater than a certain proportion.

[0056] Specifically, in region a of the current image ij The average grayscale value is greater than the first preset threshold, and the current image region a ij When the standard deviation is greater than the second preset threshold, the region a is determined. ij The image changed.

[0057] In region a of the current image ij The average grayscale value is less than or equal to the first preset threshold, or the grayscale value of region a in the current image is less than or equal to the first preset threshold. ij When the standard deviation is less than or equal to the second preset threshold, the calculation of region a of the current image continues. ij and the area m of the background grayscale image ij The absolute difference in the region a of the current image being calculated. ij and the area m of the background grayscale image ij When the absolute difference is greater than the third preset threshold, the region a of the current image is determined. ij relative to the area m of the background grayscale image ij The image changed.

[0058] Specifically, in identifying the region a in the current image where the image has changed... ij When the proportion of the number of images in a given area exceeds a certain threshold, it is determined that a change has occurred in the current image.

[0059] For specific ratios, the video app offers sensitivity options of extremely low, low, moderate, high, and extremely high, depending on the monitored scene. These represent values ​​within the range of 2% to 30%, specifically 20%, 10%, 8%, 5%, and 2%. The default is moderate, which corresponds to the area of ​​image change (a). ij The number is greater than 8% of the total number of areas. If the surveillance camera is installed outdoors facing the street, it is prone to changes. The camera sensitivity can be set to very low, and the corresponding percentage will be larger, for example, 20%. If the camera is installed indoors, the camera sensitivity can be set to very high, and the corresponding percentage will be smaller, for example, 2%.

[0060] Based on the results of the image change judgment and the linkage instructions of the linkage device, it is determined whether the current image needs to be reported for analysis. Specifically, when the background model determines that there is a change in the area or the brightness of the light in the current image, a screenshot is taken and the captured image is uploaded to the object storage.

[0061] When the background model determines that there is no change in the area or brightness of the current scene, but receives a linkage command from the linkage device, it also takes a screenshot and uploads the captured image to the object storage.

[0062] Based on the linkage command of the linkage device, the camera rotates up and down and left and right, and takes a screenshot every time it rotates a specified angle, and uploads the captured image to the object storage.

[0063] In one optional embodiment, based on the linkage command of the linkage device, the camera can rotate 0° to 360° in the horizontal direction and 90° in the vertical direction.

[0064] Optionally, the specified angle includes a rotation of 0° to 60° in the horizontal direction and a rotation of 10° to 20° in the vertical direction.

[0065] For example, the specified angle includes rotations of 15°, 20°, and 30° in the horizontal direction and rotations of 10° to 20° in the vertical direction.

[0066] In one specific implementation, for a specific analysis task of electric vehicle identification, the system judges whether the electric vehicle is moving in the electric vehicle image, whether smoke appears in the electric vehicle image while it is charging, and whether someone enters the image where the electric vehicle is placed, etc., and obtains the image change judgment result. The image change judgment result is, for example, that the electric vehicle has moved, smoke appears in the electric vehicle image, or someone enters the image where the electric vehicle is placed, etc.

[0067] When the background model determines that the area of ​​a region in the current image changes (e.g., the area of ​​a specified region becomes smaller) or the light intensity changes (e.g., the light intensity changes due to the presence of smoke), it takes a screenshot of the image and uploads the captured image to object storage.

[0068] In another specific implementation, such as Figure 3 As shown, the camera reports to the streaming media server whether it supports the smoke sensor linkage capability (corresponding to...). Figure 3 (4.1.1) When the streaming media server determines that it supports smoke sensor linkage, the video app displays the linkage switch (corresponding to...). Figure 3 (4.1.2) and the user can turn on the smoke sensor's linkage switch on the video APP's display interface (corresponding to Figure 3 (4.1.3).

[0069] When the user turns on the smoke sensor's linkage switch, the streaming media server sends the smoke sensor switch configuration (corresponding to...) to the camera. Figure 3 (4.1.4) Next, the camera sends information to the smoke sensor, such as notifying the smoke sensor to activate smoke detection (corresponding to...). Figure 3 (4.1.5 in the text). Smoke sensor (i.e., corresponding to...) Figure 3 The smoke sensor device in the middle performs smoke detection. When smoke is detected, it reports the detection information to the camera, such as whether smoke is detected or not (corresponding to...). Figure 3 (4.1.6 in the document). Next, the camera uploads the captured image to object storage (corresponding to...). Figure 3 Section 4.1.7 uploads a screenshot of the current page to object storage. Object storage controls the camera to rotate, rotating by a specified angle each time (corresponding to...). Figure 3 Take a screenshot from a specific angle (as shown in the image) and upload it to object storage (corresponding to...). Figure 3 (4.1.8 The camera at the foot of the mountain and rotating left and right takes a screenshot and uploads it to object storage at each specific rotation angle).

[0070] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.

[0071] Next, in step S103, based on the result of the image change judgment, it is determined whether to trigger the linkage device to perform the corresponding linkage detection, and the return information corresponding to the image screenshot is sent to the current monitoring camera.

[0072] In one specific implementation, when the image change judgment result includes changes in light brightness and a specified object (e.g., smoke), it is determined that the linkage device is triggered, and the return information corresponding to the image screenshot is sent to the current monitoring camera. Specifically, after the object storage successfully receives the screenshot from the current monitoring camera, it sends return information, which is, for example, the screenshot's intranet access address or access link (e.g., URL).

[0073] Next, the triggered linkage device performs corresponding linkage detection. The linkage device is a smoke sensor. The triggered smoke sensor performs smoke detection. Upon receiving a control command from the smoke sensor, the captured image is uploaded to object storage. Alternatively, the current camera uploads a screenshot of the current page every 5 seconds until it receives a message from the smoke sensor indicating the smoke has disappeared, at which point it stops uploading the captured image to object storage.

[0074] In one specific implementation, the streaming media server sends an activation notification, specifically sending the smoke sensor device linkage switch enabled by the video app to the current surveillance camera. The current surveillance camera then notifies the smoke sensor device to activate smoke detection. Alternatively, the user can activate the smoke sensor device linkage switch in the video app, triggering the smoke sensor device to perform a corresponding linkage test. When the smoke sensor device detects smoke, it reports it to the current surveillance camera, which then processes the image and uploads the captured screenshot of the current page, thereby achieving smoke detection and recognition.

[0075] It should be noted that in other embodiments, the system may also include reporting to the streaming media server whether the current surveillance camera supports the linkage capability with the smoke sensor device. Based on the data reported to the streaming media server, the video app, if the current surveillance camera supports the linkage capability with the smoke sensor device, will display the smoke sensor linkage switch. The above is merely an optional example and should not be construed as limiting the invention.

[0076] In another specific implementation, when the image change judgment result includes a person, a linkage device is triggered, such as an FTTR (Fiber to the Room) device. The FTTR device extends an optical fiber from the communication base station to the target location (e.g., a parking area for an electric vehicle, or a specific location inside a home). Because the human body absorbs and reflects wireless signals, when someone enters the wireless signal range, they absorb some signal energy and generate reflection and scattering, changing the signal propagation direction and generating an echo signal. Based on the signal propagation direction and the echo signal, the presence of a person can be determined.

[0077] Specifically, FTTR-related devices are used to send control commands, such as rotating the camera, to the current surveillance camera. After receiving the control commands, the current surveillance camera begins to rotate left and right and up and down to look for people, animals, or other objects (e.g., objects not present in the original still image). When a person appears in the searched image, a screenshot is taken and uploaded to the object storage.

[0078] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.

[0079] Next, in step S104, the return information corresponding to the screenshot, the current information of the monitoring camera, and the time of the screen change are reported to the scheduling server so that they can be scheduled to one or more corresponding analysis servers for analysis.

[0080] Specifically, the surveillance camera will report the received information (such as screenshots of intranet access addresses or access links), the camera's MAC address (the MAC address is a unique hardware identifier used to identify each surveillance camera in the network), and the time of the image change to the AI ​​scheduling server so that it can be scheduled to one or more corresponding analysis servers for analysis.

[0081] Next, when the AI ​​scheduling server receives the image change information reported by the current monitoring camera, it queries which AI packages the camera has subscribed to based on the MAC information of the current monitoring camera. If it finds that the camera has enabled the face recognition AI package and the fall detection AI package, it will schedule the camera to the algorithm analysis server with the convolutional neural network (CNN) installed to analyze the face, and to the algorithm analysis server with the recurrent neural network (RNN) installed to analyze the fall.

[0082] Specifically, when the corresponding AI algorithm analysis server receives an analysis request from the AI ​​scheduling server, the request parameters include the camera's MAC address, the screenshot time, and the URL for accessing the screenshot object storage via the intranet. First, the screenshot is downloaded from the object storage via the intranet. Then, the screenshot is scaled, cropped, and normalized to adapt to the input requirements of the relevant algorithm model. Next, key features such as edges, textures, colors, human poses, and motion patterns are extracted from the screenshot. Finally, the extracted image features are input into the relevant algorithm model to calculate and output a confidence score. The confidence score is a value between 0 and 1, representing the model's confidence level in the analysis results. For example, in fall detection analysis, if the relevant algorithm model is an RNN model, and the RNN model outputs a confidence score of 0.9, it means that the RNN model's output indicates a 90% certainty that the prediction is correct. For example, the algorithm analysis server returns the confidence score calculated by the RNN model to the algorithm scheduling server.

[0083] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.

[0084] Next, in step S105, the analysis results of one or more corresponding analysis servers are received and the analysis results are notified to the video APP. The analysis results include AI recognition analysis results, AI events, and event start times.

[0085] Specifically, the AI ​​algorithm scheduling server receives analysis results from one or more corresponding analysis servers. It then integrates and processes the recognition and analysis results from multiple AI algorithm analysis servers, for example, by notifying a video app via push notification. The notification message carries the AI ​​recognition result text, specifically including the push message title (e.g., the type of AI event detected, such as fall detection, face recognition, or smoke detection), the push message content (e.g., using "content" to represent the detailed content of the detected AI event, such as "camera name detected someone fell"), the device MAC address, the custom AI event type (e.g., object movement: 1 indicates someone is present; 2 indicates fall detection; 3 indicates smoke detection), and the start time of the recognized AI event.

[0086] In one specific implementation, when a video app is notified via push notification, the user can open the notification and be redirected to the current surveillance camera playback page of the video app. The recognized AI event will start playing from the event start time in the push notification, such as a stranger detected by face recognition, a person falling detected by fall detection, a fire detected by fire detection, and a motor vehicle / non-motor vehicle / person / pet intrusion detected by object recognition.

[0087] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.

[0088] Compared with existing technologies, this invention is based on the dynamic detection of various scene changes by a surveillance camera, such as changes in the area of ​​a designated region and changes in light brightness, and the screenshots are transmitted to the cloud AI for real-time analysis. The image changes are judged based on the constructed background model to determine whether to trigger the linkage device for corresponding linkage detection, and the return information corresponding to the image screenshot is sent to the current surveillance camera. The return information corresponding to the image screenshot, the information of the current surveillance camera, and the time of the image change are further reported to the scheduling server, so as to schedule one or more corresponding analysis servers for analysis, and the analysis results are notified to the video APP. This effectively solves the problem of slow detection delay of cameras in perceiving small changes in light brightness.

[0089] In addition, linking the camera with the FTTR device effectively solves the problem of the camera being unable to detect people, animals or other objects due to its placement angle.

[0090] In addition, deploying the AI ​​algorithm analysis server and object storage on the same intranet effectively solved the problem of bandwidth consumption from downloading a large number of images.

[0091] Example 2

[0092] The following are system embodiments of the present invention, which can be used to execute the method embodiments of the present invention. For details not disclosed in the system embodiments of the present invention, please refer to the method embodiments of the present invention.

[0093] Figure 4 This is a schematic diagram of an example of a real-time video analysis system based on the combination of dynamic detection and cloud AI according to the present invention. The following will refer to... Figure 4 The present invention describes a real-time video analysis system. This system is used to execute the real-time video analysis method described in the first aspect of the invention.

[0094] like Figure 4 As shown, the real-time video analysis system 400 includes a receiving and processing module 410, a screen change judgment module 420, a determination module 430, a scheduling module 440, and a notification module 450.

[0095] In one specific embodiment, the receiving and processing module 410 is used to receive a connection address request from the current surveillance camera. After authorization and authentication, it provides an available connection address to the current surveillance camera, enabling the current surveillance camera and the streaming media server to establish a long connection. The image change judgment module 420 is used to upload the real-time captured video clips to object storage, take screenshots based on the real-time captured video clips, and judge image changes based on the constructed background model, specifically including changes in area and light brightness. The determination module 430 determines whether to trigger the linkage device to perform corresponding linkage detection based on the image change judgment result, and sends the return information corresponding to the image screenshot to the current surveillance camera. The scheduling module 440 is used to report the return information corresponding to the image screenshot, the information of the current surveillance camera, and the time of image change to the scheduling server, so as to schedule one or more corresponding analysis servers for analysis. The notification module 450 is used to receive the analysis results of one or more corresponding analysis servers and notify the video APP of the analysis results. The analysis results include AI recognition analysis results, AI events, and event start times.

[0096] According to the optional implementation method, a background model is constructed, specifically a set of background grayscale images corresponding to each color image captured by the surveillance camera under unobstructed conditions.

[0097] To acquire multiple frames of color images from a surveillance camera under unobstructed conditions, convert these images into multiple background grayscale images. Then, calculate the average grayscale value at the same pixel location within these grayscale images using the following formula:

[0098]

[0099] Where mean(g)xy represents the average gray value of all pixel positions in multiple background grayscale images; It is the gray value of the background grayscale image corresponding to the i-th color image at pixel position (x, y), where x represents the horizontal coordinate of pixel position (x, y), y represents the vertical coordinate of pixel position (x, y), and i is a positive integer, specifically 1, 2, ..., n, where n represents the number of background grayscale images.

[0100] According to an optional implementation, the real-time captured current image and the background grayscale image are cut into regions of the same size in both the horizontal and vertical directions. For each region, the mean and standard deviation of the grayscale of the current image are calculated, and the mean and standard deviation of the grayscale of the background grayscale image are also calculated.

[0101] Based on the segmented regions, determine whether the current image has undergone a change in appearance. Specifically, determine if the proportion of regions with changed appearances in the current image to the total number of regions is greater than a certain proportion.

[0102] According to the optional implementation method, based on the judgment result of the screen change and the linkage instruction of the linkage device, it is determined whether the current screen needs to be reported for analysis. Specifically, when the background model determines that the area of ​​the current screen has changed or the brightness of the light has changed, a screenshot is taken and the captured screen is uploaded to the object storage.

[0103] When the background model determines that there is no change in the area or brightness of the current scene, but receives a linkage command from the linkage device, it also takes a screenshot and uploads the captured image to the object storage.

[0104] According to an optional implementation, based on the linkage command of the linkage device, the camera is rotated up and down and left and right, and a screenshot is taken every time it rotates by a specified angle, and the captured image is uploaded to the object storage.

[0105] In one optional embodiment, based on the linkage command of the linkage device, the camera can rotate 0° to 360° in the horizontal direction and 90° in the vertical direction. The specified angle includes rotating 0° to 60° in the horizontal direction and 10° to 20° in the vertical direction.

[0106] According to an optional implementation, the AI ​​algorithm analysis server and object storage are deployed on the same intranet. The same intranet is interconnected through network devices such as routers and switches. The devices in the same intranet usually have private identification addresses, so that data transmission only takes place on the intranet and does not require an external network.

[0107] It should be noted that, due to Figure 4 The method performed by the device and Figure 1 The methods in the examples are largely the same, therefore, the descriptions of the same parts have been omitted.

[0108] Compared with existing technologies, this invention is based on the dynamic detection of various scene changes by a surveillance camera, such as changes in the area of ​​a designated region and changes in light brightness, and the screenshots are transmitted to the cloud AI for real-time analysis. The image changes are judged based on the constructed background model to determine whether to trigger the linkage device for corresponding linkage detection, and the return information corresponding to the image screenshot is sent to the current surveillance camera. The return information corresponding to the image screenshot, the information of the current surveillance camera, and the time of the image change are further reported to the scheduling server, so as to schedule one or more corresponding analysis servers for analysis, and the analysis results are notified to the video APP. This effectively solves the problem of slow detection delay of cameras in perceiving small changes in light brightness.

[0109] In addition, linking the camera with the FTTR device effectively solves the problem of the camera being unable to detect people, animals or other objects due to its placement angle.

[0110] In addition, deploying the AI ​​algorithm analysis server and object storage on the same intranet effectively solved the problem of bandwidth consumption from downloading a large number of images.

[0111] Figure 5 This is a schematic diagram of an embodiment of an electronic device according to the present invention.

[0112] like Figure 5 As shown, the electronic device is embodied in the form of a general-purpose computing device. There can be one or more processors working collaboratively. This invention also does not preclude distributed processing, meaning that processors can be distributed across different physical devices. The electronic device of this invention is not limited to a single entity, but can also be the sum of multiple physical devices.

[0113] The memory stores a computer-executable program, typically machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some steps of the method.

[0114] The memory includes volatile memory, such as random access memory (RAM) and / or cache memory, and may also be non-volatile memory, such as read-only memory (ROM).

[0115] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with external devices. The I / O interface can represent one or more of several bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0116] It should be understood that Figure 5 The electronic device shown is merely one example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as displays, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. Any electronic device capable of executing a computer-readable program in memory to implement the method of the present invention or at least some steps of the method can be considered as an electronic device covered by the present invention.

[0117] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software, or by combining software with necessary hardware. Therefore, as... Figure 6 As shown, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) or on a network, and includes several commands to cause a computing device (such as a personal computer, server, or network device, etc.) to execute the above-described method according to the embodiments of the present invention.

[0118] The software product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0119] The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in connection with a command execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0120] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0121] The aforementioned computer-readable medium carries one or more programs, which, when executed by a device, enable the computer-readable medium to implement the data interaction method of this disclosure.

[0122] Those skilled in the art will understand that the above modules can be distributed in the device as described in the embodiments, or they can be modified accordingly and placed in one or more devices that are unique to this embodiment. The modules in the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0123] Through the description of the above embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions of the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) or on a network, including several commands to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of the present invention.

[0124] It should be noted that the above detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0125] In the detailed description above, reference has been made to the accompanying drawings, which form part of this document. In the drawings, similar symbols typically identify similar parts unless the context otherwise indicates otherwise. The illustrated embodiments described in the detailed specification, drawings, and claims are not intended to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0126] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.< / token>

Claims

1. A real-time video analysis method based on dynamic detection and cloud AI, characterized in that, The real-time video analysis method includes: Receive a request for a connection address from the current surveillance camera. After authorization and authentication, provide the current surveillance camera with an available connection address so that the current surveillance camera and the streaming media server can establish a long connection. The real-time captured video clips are uploaded to object storage. Screenshots are taken based on the captured video clips, and image changes are determined according to a constructed background model. Based on the image change determination result, it is determined whether to trigger a linkage device for corresponding linkage detection, and the return information corresponding to the screenshot is sent to the current monitoring camera. The image change determination based on the constructed background model includes: Construct a background model, specifically a set of background grayscale images corresponding to each color image captured by the surveillance camera under unobstructed conditions; To acquire multiple frames of color images from a surveillance camera under unobstructed conditions, convert these images into multiple background grayscale images. Then, calculate the average grayscale value at the same pixel location within these grayscale images using the following formula: , Where mean(g)xy represents the average gray value at the same pixel position (x,y) in multiple background grayscale images; It is the gray value of the background grayscale image corresponding to the i-th color image at pixel position (x, y), where x represents the horizontal coordinate of pixel position (x, y), y represents the vertical coordinate of pixel position (x, y), and i is a positive integer, specifically 1, 2, ..., n, where n represents the number of background grayscale images; The real-time captured image and the background grayscale image are cut into regions of the same size in both the horizontal and vertical directions. For each region, the mean and standard deviation of the grayscale value of the current image are calculated, as well as the mean and standard deviation of the grayscale value of the background grayscale image. Based on the segmented regions, determine whether the current image has undergone a change in appearance. Specifically, determine that the current image has undergone a change in appearance when the proportion of the number of regions in the current image that have undergone a change in appearance to the total number of regions is greater than a certain proportion. Based on the results of the image change judgment and the linkage instructions of the linkage device, it is determined whether the current image needs to be reported for analysis. Specifically, when it is determined from the background model that the area of ​​the current image has changed or the brightness of the light has changed, a screenshot is taken and the captured image is uploaded to the object storage. When the background model determines that there is no change in the area or brightness of the current image, However, when a linkage command is received from the linked device, a screenshot is also taken and the captured image is uploaded to the object storage. The returned information corresponding to the screenshot, the current information of the monitoring camera, and the time of the image change are reported to the scheduling server so that they can be scheduled to one or more AI algorithm analysis servers for analysis. The system receives analysis results from one or more AI algorithm analysis servers and notifies the video app of these results. The analysis results include AI recognition analysis results, AI events, and event start times. The AI ​​algorithm analysis servers and object storage are deployed on the same intranet, and interconnected via routers and switches. All devices on the same intranet have private identifier addresses, ensuring that data transmission occurs only within the intranet without the need for an external network.

2. The real-time video analysis method based on dynamic detection and cloud AI as described in claim 1, characterized in that, Further includes: Based on the linkage command of the linkage device, the camera rotates up and down and left and right, and takes a screenshot every time it rotates a specified angle, and uploads the captured image to the object storage.

3. The real-time video analysis method based on dynamic detection and cloud AI as described in claim 2, characterized in that, Further includes: Based on the linkage command of the linkage device, the camera can rotate 0° to 360° in the horizontal direction and 90° in the vertical direction. The specified angle includes rotating 0° to 60° in the horizontal direction and 10° to 20° in the vertical direction.

4. A real-time video analysis system based on dynamic detection and cloud-based AI, characterized in that, The system implements the real-time video analysis method based on dynamic detection and cloud AI as described in any one of claims 1 to 3, wherein the real-time video analysis system comprises: The receiving and processing module is used to receive the application request for the connection address from the current surveillance camera. After authorization and authentication, it provides the current surveillance camera with an available connection address so that the current surveillance camera and the streaming media server can establish a long connection. The image change judgment module is used to upload the real-time captured video clips to object storage, take screenshots based on the real-time captured video clips, and judge image changes based on the constructed background model, specifically including changes in area and changes in light brightness. The determination module, based on the judgment result of the screen change, determines whether to trigger the linkage device to perform the corresponding linkage detection, and sends the return information corresponding to the screen screenshot to the current monitoring camera; The scheduling module is used to report the return information corresponding to the screenshot, the current information of the monitoring camera, and the time of the image change to the scheduling server, so as to schedule it to one or more AI algorithm analysis servers for analysis; The notification module is used to receive the analysis results from one or more AI algorithm analysis servers and notify the video APP of the analysis results. The analysis results include AI recognition analysis results, AI events, and event start times.

Citation Information

Patent Citations

  • AI detection method based on cloud algorithm

    CN111354154A