Video linkage method
By adopting segmented storage and event serialization transmission strategies in video linkage technology, the camera generates audio and video files and stores them after receiving event information, solving the problem of high storage and network bandwidth requirements in the prior art, and achieving more efficient video system deployment and transmission efficiency.
Patent Information
- Application Number
- CN202510267126.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
AI Technical Summary
In the era of Internet of Things with massive video equipment access, existing video linkage technology has problems such as high storage and network bandwidth requirements, poor performance and fluency, which seriously restricts the large-scale deployment of video systems.
A video linkage method is adopted. After receiving event information, the camera generates audio and video files and stores them every several times. The event information is separated from the file and transmitted. The platform dynamically synthesizes the video files corresponding to the linkage event required based on the trigger event.
Through segmented storage and event serialization transmission strategies, the storage space usage and network bandwidth requirements are significantly reduced, and the transmission efficiency and the scalability of the video system are improved.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video linkage, and particularly relates to a video linkage method. Background Art
[0002] In the video field, the video linkage technology based on event triggering has been widely applied. Traditional technical solutions usually adopt a direct mapping mechanism of "event - video", that is, when a specific event is detected, the associated camera is triggered to start shooting, and the complete video is uploaded to the platform. Another relatively common method is that the platform or NVR device receives the audio - video stream and event information in real time, and the audio - video stream is intercepted into individual event - related material files in real time on the platform or NVR. Both of these two solutions have relatively high requirements for storage and network bandwidth, and there will be many performance and fluency problems in the actual application process.
[0003] The superposition effect of these technical defects is becoming increasingly prominent in the context of the massive access of video devices in the Internet of Things era, seriously restricting the large - scale deployment and implementation of the video system. Therefore, there is an urgent need for a video linkage method with redundancy removal and dynamic video synthesis capabilities to break through the bottleneck of the existing technical framework. Summary of the Invention
[0004] The purpose of the present invention is to provide a video linkage method to overcome the deficiencies in the prior art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions: The present application discloses a video linkage method, which specifically includes the following steps: S1. The camera works, continuously generating a video stream and an audio stream; when receiving event information, it enters step S2; S2. After the camera receives the event information, it parses the event information; S3. According to the parsed event information, the camera generates and stores an audio - video file every certain period of time; S4. The camera transmits the event information and the stored audio - video files to the platform in sequence; S5. After receiving the event information and the audio - video files, the platform dynamically and real - time synthesizes the video file corresponding to the required linkage event based on the trigger event.
[0006] Preferably, the event information includes a trigger device ID, a linkage device ID, a camera ID, a capture start time, a capture end time, and an event current time.
[0007] Preferably, in step S3, the camera takes the capture start time as the starting point and the capture end time as the ending point, and generates and stores an audio - video file every 10 seconds.
[0008] Preferably, the name of the audio - video file is camera device ID - sensor ID - event start timestamp - event end timestamp - current file start time - file type - frame rate - event sequence number.
[0009] Preferably, step S5 specifically includes the following sub - steps: S51. Obtain the camera ID, capture start time, and capture end time according to the event information, and organize them. S52. Obtain the corresponding camera ID, start time, end time, frame rate, and sequence number information of the audio - video file according to the audio - video file name, and organize them. S53. According to the required event information, find the audio - video file corresponding to the capture start time and capture end time. The specific operation is as follows: find the sequence of the audio - video file where the capture start time TEstart is greater than the event start time TMstart and the capture end time TEend is greater than the time end time TMend. Use (TEstart – TMstart) / T as the start sequence number and (TEstart – TMstart) / T+(TEend - TEstart) / T as the end sequence number. S54. According to the frame rate information, find the specific start frame and end frame information. The specific operation is as follows: (TEstart – TMstart)%T is the corresponding start frame time, and (TEend - TEstart)%T is the end frame time. S55. Find the corresponding start frame, end frame, and the complete file sequence range in the audio - video file, and merge the video file and the audio file into a playable media file.
[0010] Preferably, the playable media files in step S55 include media files in MP4, MOV, and MKV formats.
[0011] Preferably, in step S55, according to user needs, it can be merged into one or more audio files. The file name format of the audio file is: camera ID - triggering device ID - associated device ID - triggering time.
[0012] Advantages of the present invention: Segmented storage mechanism Through the time - window slicing technology, independent audio - video files are generated every certain period of time, realizing dynamic storage granularity control. Compared with the traditional continuous recording mode, the storage space occupation is significantly reduced.
[0013] (2) Precise transmission control Adopt the event serialization transmission strategy (sorting by the timestamp and serial number in the file name), combined with the metadata pre - transmission mechanism (separately transmitting event information and files), to improve the network bandwidth utilization rate.
[0014] (3)Resume interrupted transfer guarantee Based on the timestamp - serial number indexing system of the file name, after the network is interrupted, only the missing fragments need to be re - transmitted instead of the entire video, which greatly improves the transmission efficiency.
[0015] Frame - level positioning technology Through the remainder calculation mechanism of (TEstart–TMstart)%T, accurately locate the start frame and end frame.
[0016] Intelligent synthesis engine Support multi - format output (MP4 / MOV / MKV) and multi - audio - track mixing (multiple audio channels can be combined) to meet the requirements of different business scenarios.
[0017] Fault - tolerant recovery mechanism When a certain segment of the file is damaged, the system can perform interpolation repair through the timestamp and frame rate information of adjacent segments, and the video availability is significantly improved compared with traditional solutions.
[0018] The features and advantages of the present invention will be described in detail through embodiments. Specific implementation manners
[0019] To make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail through embodiments below. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the scope of the present invention. In addition, in the following description, the descriptions of well - known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0020] A video linkage method of the present invention specifically includes the following steps: S1. The camera works, continuously generating video streams and audio streams; when receiving event information, enter step S2; S2. After the camera receives the event information, parse the event information; S3. According to the parsed event information, the camera generates and stores audio - video files at regular intervals; S4. The camera transmits the event information and the stored audio - video files to the platform in sequence; S5. After the platform receives the event information and the audio - video files, dynamically and real - time synthesize the corresponding video files of the required linkage events based on the triggering events.
[0021] Specifically, the video file transmission and event separation mechanism: 1: The camera linkage event supports logic such as card swiping, location, area entry / exit, and remote control. This event contains event source information (such as the ID information of wearable devices, identity information of face recognition), scene information (such as area location, attributes, location markers), device information (the ID information of the camera itself), and time information (start time, end time, duration, etc.). This event information is processed separately and sent to the server for storage and management.
[0022] 2: The camera operates in a listening state, continuously generating video and audio streams. The camera stores at least 10 seconds of historical audio-visual data streams. The actual stored real length is calculated based on the frame rate. If the current audio-visual material is 30 frames per second, the actual size of each stored file is set to 10 * 30 = 300 frames.
[0023] 3: The audio-visual data streams of the camera support two usage forms. One is for the real-time live broadcast function, and the data frames of the audio-visual are directly transmitted using librtmp and sent to the video server. The other is to parse the event information and perform local storage after receiving the event. The following decomposes the event linkage function.
[0024] 4: After the camera receives the event information, it parses the event information to understand the form of video capture, the capture time point (current time, how many seconds before the current time, how many seconds after the current time, etc.), and the video end time (how long it lasts based on the start time, waiting for the capture end event, etc.). Starting from the start time, the camera begins to store the audio-visual stream starting from the corresponding time point. An audio-visual file is generated locally every ten seconds. After all events are over, the local storage of audio-visual files stops. The files are generated according to a fixed duration. If the end event corresponds to a time less than the required file time length, they are stored according to the fixed file duration.
[0025] 5: The camera transmits the event information, each video, and audio file segment to the platform in sequence.
[0026] Definition of event information format: Trigger device ID, linked device ID, camera ID, capture start time, capture end time (or fixed duration, or timeout time), event current time. Among them, if there is no capture end time and it is 0, it means only starting the capture start event. If the capture start time is 0, it means only stopping the capture end event. If there is only an event start event in the event and no end event, it is calculated according to the timeout time.
[0027] Definition of audio-visual file name: Camera device ID - sensorID - start timestamp of this group of events - end timestamp of this group of events - start time of the current file - file type - frame rate - event sequence number Among them: sensorID is used to distinguish the definitions of the sensor corresponding to the directly connected video stream on the camera and the audio input mic. The file types are VD for video stream files, AD for audio stream files, and JPG for picture files. When the file is not the last file of the group of events, the event end timestamp is 0, and the event sequence number starts from 0 and accumulates to the last file. Among them, the start timestamp of the audio-visual material can be set to be about 10 seconds (configurable) ahead of the current time. That is, the camera can store 10 seconds (configurable) of historical data.
[0028] 6: After the platform receives the events and independent files, as needed, it dynamically and real-time synthesizes the corresponding video files of the required linked events based on the triggered events.
[0029] A: Organize the event information, obtain the camera ID, the capture start time TEstart, and the capture end time information TEend. The time information is accurate to the second. B: Organize the audio-visual material files. According to the file name information, obtain the camera ID, the event start time TMstart, the event end time TMend, the frame rate, and the sequence number information corresponding to the material files.
[0030] C: Since the length time of the audio-visual material file is fixed at T (for example, 10 seconds), find the sequence of the audio-visual material corresponding to the start time and end time of the event, and find the material corresponding to TEstart > TMstart and TEend > TMend. (TEstart – TMstart) / T is the start sequence number, and (TEstart – TMstart) / T + (TEend - TEstart) / T is the end sequence number.
[0031] D: For the corresponding material file, according to the frame rate information, find the specific start frame and end frame information. (TEstart – TMstart)%T is the start frame time corresponding to the start sequence material, and (TEend - TEstart)%T is the end frame time in the end sequence material. According to the value of the frame rate, find the corresponding start and end frame data in the corresponding material segment.
[0032] E: Based on the above steps, respectively find the corresponding start frame, end frame, and the complete material sequence range in the video material and the audio material, and merge the video material and the audio material into a playable media file, such as media files in formats such as MP4, MOV, and MKV. During the merging process, according to the user's needs, one or multiple audio materials can be merged and presented in the form of audio tracks or other forms.
[0033] F: Name and store the audio file, which corresponds to the linked video file at this time. The file name format is: Camera ID - Trigger Device ID - Linked Device ID - Trigger Time. Embodiment
[0034] A certain smart kindergarten deploys multiple types of trigger devices (infrared fence sensors, abnormal sound detectors, fall recognition) in dangerous areas (playground areas, pool sides, kitchen entrances). In traditional solutions, when multiple sensors are triggered simultaneously (such as a child climbing over the fence triggering the infrared sensor + fall recognition, each device generates a monitoring video independently, resulting in multiple duplicate records.
[0035] (1) Embodiment of the present invention: During the lunch break activity, a child wearing a device triggers simultaneously in the play area: an infrared fence sensor, detecting an out-of-bounds behavior (14:15:00 - 14:16:00); an abnormal sound detector, identifying the sound of a fall (14:15:20 - 14:16:20).
[0036] (2) Technical implementation process: a. The trigger device generates events and synchronizes with the camera, Event 1 (start = 14:15:00, end = 14:16:00); Event 2 (start = 14:15:20, end = 14:16:20); b. The camera parses the event information; c. According to the parsed event information, the camera generates and stores audio - video files every 10 seconds; d. The camera transmits the event information and the stored audio - video files to the platform in sequence; (3) Implementation results: Main video panoramic view (14:15:00 - 14:16:20), key - frame screenshots (moment of out - of - bounds + fall action), safety event rating report (based on duration and risk level) Enhanced event visualization: Automatically generate a dual - event timeline: mark the out - of - bounds period (14:15:00 - 16:00) in red and the fall - risk period (14:15:20 - 16:20) in yellow; overlay a safety prompt layer: insert virtual warning boxes (such as "Out - of - bounds area!" pop - up windows) and voice prompts ("Fall risk detected") in the video (4) Comparison of technical effects: Index Traditional solution The present invention Optimization rate Number of video files 2 groups × 6 shards = 12 8 33% Storage space 12 × 10MB = 120MB 8 × 10MB = 80MB 33% Network transmission time 1 minute and 12 seconds 44 seconds 39% (5) Verification with measured data: During the three-month deployment in a provincial demonstration kindergarten, the average storage volume of dual-device trigger events decreased from 1.2 GB to 640 MB; the time taken for event review was shortened from 8 minutes per time to 2.5 minutes per time; and the complaint rate of parents due to untimely video reception decreased by 65%.
[0037] This embodiment proves that the present invention has achieved a breakthrough of "multi-risk perception - single video traceability" in the children's guardianship scenario. While ensuring security, it significantly reduces data redundancy and provides a compliant and efficient intelligent guardianship solution for the education industry.
[0038] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, or improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A video linkage method, characterized in that: The specific steps include: S1, the camera works and continuously generates video stream and audio stream; after receiving event information, it goes to step S2; S2, after receiving the event information, the camera analyzes the event information; S3. Based on the parsed event information, the camera generates and stores audio and video files at certain intervals; S4, the camera transmits the event information and the stored audio and video files to the platform in order; S5. After receiving the event information and the audio and video files, the platform dynamically and in real time synthesizes the video files corresponding to the required linkage events based on the triggering events.
2. A video linkage method as claimed in claim 1, characterized in that: The event information includes trigger device ID, linkage device ID, camera ID, capture start time, capture end time and event current time.
3. A video linkage method as claimed in claim 2, characterized in that: In step S3, the camera takes the capture start time as the starting point and the capture end time as the end point, and generates and stores an audio and video file every 10 seconds.
4. A video linkage method as claimed in claim 1, characterized in that: The name of the audio and video file is camera device ID-sensorID-event start timestamp-event end timestamp-current file start time-file type-frame rate-event sequence number.
5. A video linkage method as claimed in claim 1, characterized in that: Step S5 specifically includes the following sub-steps: S51, according to the event information, obtain the camera ID, capture start time, capture end time, and organize them; S52, according to the audio and video file name, obtain the camera ID, time start time, time end time, frame rate and serial number information corresponding to the audio and video file, and organize them; S53, according to the required event information, search for the audio and video files corresponding to the corresponding capture start time and capture end time; the specific operation is as follows: find the sequence of audio and video files corresponding to the capture start time TEstart greater than the event start time TMstart and the capture end time TEend greater than the time end time TMend, with (TEstart-TMstart) / T as the start sequence number and (TEstart-TMstart) / T+(TEend-TEstart) / T as the end sequence number; S54, according to the frame rate information, find the specific start frame and end frame information, the specific operation is as follows: (TEstart-TMstart)%T is the corresponding start frame time, (TEend-TEstart)%T is the end frame time; S55: Find the corresponding start frame, end frame, and complete file sequence range in the audio and video files, and merge the video file and the audio file into a playable media file.
6. A video linkage method as claimed in claim 5, characterized in that: The media files that can be played in step S55 include media files in MP4, MOV, and MKV formats.
7. A video linkage method as claimed in claim 5, characterized in that: In step S55, the audio files can be merged into one or more audio files according to user needs; the file name format of the audio file is: camera ID-trigger device ID-linked device ID-trigger time.