Alarm positioning method and system based on 5G video multi-modal large model analysis
The alarm location method based on 5G video multimodal large model analysis automatically obtains the alarm location using video data from citizens' mobile phone cameras, solving the problems of low positioning efficiency, poor accuracy and high cost in existing technologies, and achieving efficient and convenient alarm location.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XINZHI DAOSHU (SHANGHAI) TECH CO LTD
- Filing Date
- 2025-08-25
- Publication Date
- 2026-06-23
AI Technical Summary
The existing 110 emergency call system relies on manual inquiry, carrier location, or SMS interaction to obtain the location of alarms, resulting in low location efficiency, poor accuracy, and high cost. It cannot achieve automatic location without relying on third-party platforms or in a non-interactive manner.
By using 5G video multimodal big data analysis, real-time video footage from citizens' mobile phone cameras is acquired, preprocessed, and keyframes are extracted to extract geographic information, spatial building information, and key graphic and textual information. Combined with dynamic scene content, surveillance cameras are selected and alarm locations are determined.
It enables automatic alarm positioning without relying on third-party platforms or engaging in non-interactive operations, improving positioning accuracy and convenience while reducing dependence on operators and costs.
Smart Images

Figure CN121214286B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image processing and positioning, and in particular to an alarm positioning method and system based on 5G video multimodal large model analysis. Background Technology
[0002] The 110 emergency call system is an information system used by public security organs to quickly respond to public alarms, dispatch police resources, and carry out on-site handling. It has functions such as receiving alarms, dispatching police, command and dispatch, and closed-loop management of police situation tracking.
[0003] In existing technologies, the 110 emergency call system needs to obtain the current location of the caller during the call process. Traditional methods for obtaining location include: relying on the dispatcher to question the caller, which is time-consuming, depends on the communication skills of both parties, and is highly dependent on the caller's ability to articulate their situation. Obtaining location through mobile operators is problematic, as operators are often unwilling to open their interfaces, and the provided locations are mostly based on cell tower positioning, resulting in poor accuracy. Furthermore, it requires interfacing with multiple operators, leading to high overall costs. Location can be obtained through location networks, but this requires significant annual investment in location service fees. SMS-assisted alarm location systems require the dispatcher to send a link to the caller, who must click and agree. If there is no network connection or the caller does not agree, location cannot be determined. Therefore, how to automatically complete alarm location without relying on third-party service platforms and using non-interactive methods has become a pressing issue. Summary of the Invention
[0004] The purpose of this invention is to provide an alarm location method and system based on 5G video multimodal large model analysis to solve the problems mentioned in the background art.
[0005] In a first aspect, this application provides an alarm location method based on 5G video multimodal large model analysis, the method comprising:
[0006] When a citizen uses a 5G mobile phone to report an emergency, the system acquires real-time video footage from the citizen's mobile phone camera and preprocesses the real-time video footage to obtain the target video footage.
[0007] Keyframes are extracted from the target video frame to obtain multiple keyframe frames. Geographic information, spatial architectural information, and key graphic information are extracted from the keyframe frames to generate a spatial semantic information set.
[0008] Extract dynamic scene content from the target video frame, obtain the dynamic scene subject and scene subject action based on the dynamic scene content, and obtain dynamic semantic information based on the dynamic scene subject and scene subject action;
[0009] Based on the spatial semantic information set, a range query is performed on the preset positioning map to obtain the target positioning range;
[0010] Capture the comparison reference images captured by the surveillance cameras within the target positioning range, compare the spatial semantic information set and the dynamic semantic information with the comparison reference images to obtain the comparison similarity;
[0011] Multiple surveillance cameras are filtered based on the comparison similarity to obtain target cameras. The camera location information of the target cameras is extracted, and the alarm location is determined based on the camera location information.
[0012] Preferably, the step of acquiring real-time video footage from a citizen's mobile phone camera and preprocessing the real-time video footage to obtain the target video footage is as follows:
[0013] When a citizen uses a 5G mobile phone to report an emergency, the system automatically captures the real-time video feed from the citizen's mobile phone camera.
[0014] Extract the video resolution, video brightness, and video contrast of the real-time video frame, and obtain the real-time image quality of the real-time video frame based on the video resolution, video brightness, and video contrast.
[0015] Determine whether the real-time image quality exceeds a preset quality threshold; if it is determined that the real-time image quality does not exceed the quality threshold.
[0016] The video resolution, video brightness, and video contrast are then automatically adjusted to obtain a target video image whose real-time image quality exceeds the quality threshold.
[0017] Preferably, the step of extracting keyframes from the target video frame to obtain multiple keyframe frames, and extracting geographic information, spatial architectural information, and key graphic information from the keyframe frames to generate a spatial semantic information set, specifically includes:
[0018] Frame similarity recognition is performed on the target video frame to obtain the similarity value between each frame;
[0019] Frame recognition is performed on the image frames based on the similarity values to determine the key frame positions. Key frame extraction is then performed based on the key frame positions to obtain multiple key frame images.
[0020] The objects in the keyframe images are identified to obtain geographic information, spatial architectural information, and key graphic information, respectively.
[0021] The geographic information, spatial architectural information, and key graphic information are classified, integrated, and stored together to generate a spatial semantic information set.
[0022] Preferably, the steps of extracting dynamic scene content from the target video frame, obtaining the dynamic scene subject and scene subject action based on the dynamic scene content, and obtaining dynamic semantic information based on the dynamic scene subject and scene subject action are as follows:
[0023] Based on the target video frame, the target video frame is dynamically identified to obtain the dynamic scene content in the target video frame;
[0024] Dynamic activity recognition is performed on the dynamic scene content to obtain multiple dynamic scene subjects in the dynamic scene content;
[0025] The dynamic scene subject is continuously monitored, and the scene subject's actions during the movement of the dynamic scene subject are captured;
[0026] Static feature recognition is performed on the subject of the dynamic scene to obtain the subject's static features. Based on the subject's static features and the subject's actions in the scene, dynamic feature recognition is performed to obtain dynamic semantic information.
[0027] Preferably, the step of performing a range query on a preset positioning map based on the spatial semantic information set to obtain the target positioning range specifically includes:
[0028] Based on the geographic information in the spatial semantic information set, a large-scale positioning pointer is generated; based on the spatial building information, a medium-scale positioning pointer is generated; and based on the key graphic information, a small-scale positioning pointer is generated.
[0029] Based on the large-scale positioning pointer, the initial geographic information is matched on the preset positioning map to obtain the target's large-scale location;
[0030] Based on the mid-range pointer, the spatial building information is matched at the target's large-range location to obtain the target's mid-range location;
[0031] Based on the small-range pointer, the key graphic information is matched at the range position in the target to obtain the target positioning range.
[0032] Preferably, the step of capturing comparison reference images taken by surveillance cameras within the target positioning range, comparing the spatial semantic information set and the dynamic semantic information with the comparison reference images to obtain the comparison similarity is as follows:
[0033] Based on the target positioning range, extract data from all surveillance cameras within the target positioning range, and capture a comparative reference image within the target positioning range based on the surveillance camera data;
[0034] Substitute the geographic information, the spatial architectural information, and the key graphic information from the spatial semantic information set into the comparison reference image for feature matching to obtain a first matching value;
[0035] Based on the dynamic semantic information, video segments are searched in the comparison reference frame to obtain the target comparison video segment, and the dynamic semantic information is matched with the frequency band of the target comparison video to obtain a second matching value;
[0036] The first matching value and the second matching value are combined to generate a comprehensive matching value. The similarity is then evaluated based on the comprehensive matching value to obtain the comparative similarity.
[0037] Preferably, the steps of filtering multiple surveillance cameras based on the comparison similarity to obtain target cameras, extracting the camera location information of the target cameras, and determining the alarm location based on the camera location information are as follows:
[0038] Based on the comparison similarity, multiple surveillance cameras are filtered to obtain multiple target cameras that are closest to the crime scene;
[0039] Extract the location information, installation parameters, and camera images of multiple target cameras;
[0040] The shooting direction of the target camera is obtained based on the installation parameters, and the relative distance and relative direction between the crime scene and the target camera are obtained based on the camera image.
[0041] The relative direction between the crime scene and the target camera is determined based on the relative direction and the shooting direction.
[0042] The alarm location is determined based on the relative direction and relative distance of the target.
[0043] Preferably, after the step of filtering multiple surveillance cameras based on the comparison similarity to obtain the target camera closest to the crime scene, the method further includes:
[0044] Extract the monitoring footage from the target camera and perform a matching deviation assessment on the monitoring footage to obtain a deviation assessment value;
[0045] If the deviation evaluation value is determined to be greater than the preset evaluation threshold, then the camera resolution and camera position of multiple surveillance cameras are extracted.
[0046] The importance of the camera positions is evaluated to obtain a positional importance value for each target camera;
[0047] A weighted average is calculated based on the location importance value and the camera resolution to obtain the image credibility value of each target camera. The target cameras are then further filtered based on the image credibility value.
[0048] Preferably, after determining the alarm location based on the target's relative direction and relative distance, the method further includes:
[0049] The system acquires alarm voice information uploaded by citizens during the alarm process, transcribes the alarm voice information, and extracts location-related sentences.
[0050] Semantic recognition is performed on the location-related statements to obtain directional description information, location description information, and environmental description information, which are then combined to generate a voice location information set;
[0051] The weighted set of voice location information is sent to the set of spatial semantic information, and the set of spatial semantic information is filtered and updated.
[0052] Based on the voice location information set, the alarm location is matched for location confirmation to obtain a confirmation matching value, and it is determined whether the confirmation matching value exceeds a preset confirmation matching threshold.
[0053] If it is determined that the confirmed matching value does not exceed the confirmed matching threshold, then the monitoring camera and the camera image are re-identified according to the voice location information set to obtain the identification result, and the alarm location is updated according to the identification result.
[0054] Secondly, this application provides an alarm positioning system based on 5G video multimodal large model analysis, the system comprising:
[0055] Video capture module: When a citizen uses a 5G mobile phone to report an emergency, it acquires the real-time video footage from the citizen's mobile phone camera and preprocesses the real-time video footage to obtain the target video footage.
[0056] Static feature module: used to extract keyframes from the target video frame to obtain multiple keyframe frames, extract geographical information, spatial building information and key graphic information from the keyframe frames, and generate a spatial semantic information set;
[0057] Dynamic feature module: used to extract dynamic scene content from the target video frame, obtain dynamic scene subject and scene subject action based on the dynamic scene content, and obtain dynamic semantic information based on the dynamic scene subject and scene subject action;
[0058] Initial positioning module: used to perform range query on a preset positioning map based on the spatial semantic information set to obtain the target positioning range;
[0059] Image comparison module: used to capture comparison reference images captured by surveillance cameras within the target positioning range, compare the spatial semantic information set and the dynamic semantic information with the comparison reference images to obtain the comparison similarity;
[0060] Location determination module: used to filter multiple surveillance cameras according to the comparison similarity to obtain the target camera, extract the camera location information of the target camera, and determine the alarm location based on the camera location information.
[0061] In summary, this application includes at least one of the following beneficial technical effects:
[0062] By extracting real-time video footage captured by a citizen's 5G phone camera when using the alarm, the target video footage is obtained after preprocessing. Keyframes are extracted from the target video footage, and geographic, architectural, and key textual information is extracted from these keyframes to generate a spatial semantic information set. Moving dynamic scene content is then extracted from the target video footage, and the main subject and its corresponding actions are identified. Feature information from both is extracted to obtain dynamic semantic information. The area is then gradually delineated on a pre-defined location map based on the geographic, architectural, and key textual information to obtain the target location range. Comparison reference footage from surveillance cameras within the target location range is then extracted, and the dynamic semantic information is compared and matched with each comparison reference footage to obtain a similarity score. Surveillance cameras are then filtered based on the similarity score to identify the target cameras. Finally, the camera location information of the target cameras is extracted, and combined with the camera's installation information and camera footage, the final location is determined to pinpoint the alarm location. Finally, the voice information uploaded by citizens is identified and extracted to obtain a voice location information set containing directional, positional, and environmental descriptions. The spatial voice information set is then updated based on this voice location information set, and the accuracy of the alarm location is determined. If inaccurate, the alarm location is updated based on the voice location information set. This reduces the reliance on other operators for locating alarm callers and improves the non-interactive and convenient nature of the location process. Attached Figure Description
[0063] Figure 1 This is a flowchart of the alarm location method based on 5G video multimodal large model analysis provided in the embodiments of this application;
[0064] Figure 2 This is a block diagram of an alarm positioning system based on 5G video multimodal large model analysis provided in an embodiment of this application.
[0065] Explanation of reference numerals in the attached diagram: 1. Image capture module; 2. Static feature module; 3. Dynamic feature module; 4. Initial positioning module; 5. Image comparison module; 6. Positioning determination module. Detailed Implementation
[0066] The following is in conjunction with the appendix Figures 1-2 This application will be described in further detail, but the embodiments of the present invention are not limited thereto.
[0067] This application discloses an alarm location method and system based on 5G video multimodal large model analysis.
[0068] In this embodiment, an alarm location method based on 5G video multimodal large model analysis is described, which includes:
[0069] S100: When a citizen uses a 5G mobile phone to report an emergency, the system acquires the real-time video feed from the citizen's mobile phone camera and preprocesses the real-time video feed to obtain the target video feed.
[0070] S200: Extract keyframes from the target video frame to obtain multiple keyframe frames. Extract geographic information, spatial architectural information and key graphic information from the keyframe frames to generate a spatial semantic information set.
[0071] S300: Extract dynamic scene content from the target video frame, obtain the dynamic scene subject and scene subject action based on the dynamic scene content, and obtain dynamic semantic information based on the dynamic scene subject and scene subject action;
[0072] S400: Based on the spatial semantic information set, perform a range query on the preset positioning map to obtain the target positioning range;
[0073] S500: Captures comparison reference images captured by surveillance cameras within the target positioning range, compares the spatial semantic information set and dynamic semantic information with the comparison reference images, and obtains the comparison similarity;
[0074] S600: Filters multiple surveillance cameras based on similarity to obtain the target camera, extracts the camera location information of the target camera, and determines the alarm location based on the camera location information.
[0075] It should be noted that the above process is only the basic steps of this embodiment. In the specific implementation process, some steps may be added, reduced or modified appropriately without affecting the overall implementation effect.
[0076] The steps for acquiring real-time video footage from a citizen's mobile phone camera and preprocessing the real-time video footage to obtain the target video footage are as follows:
[0077] When a citizen uses a 5G mobile phone to report an emergency, the system automatically captures the real-time video feed from the citizen's mobile phone camera.
[0078] Extract the video resolution, video brightness, and video contrast of the real-time video frame, and obtain the real-time image quality of the real-time video frame based on the video resolution, video brightness, and video contrast.
[0079] Determine whether the real-time image quality exceeds a preset quality threshold. If the real-time image quality does not exceed the quality threshold;
[0080] It automatically adjusts the video resolution, brightness, and contrast to obtain a target video image whose real-time image quality exceeds the quality threshold.
[0081] In practice, an example is taken where a citizen in a city uses a 5G mobile phone to call the police in an emergency at a park lake. When the citizen dials 110, the system automatically activates the phone's camera to capture real-time video footage. The video resolution is 1280×720 pixels, the brightness is 80 nits, and the contrast is 40%. After extracting these parameters, the system calculates the real-time image quality. The preset quality threshold is 75 points (a comprehensive score based on resolution, brightness, and contrast). The calculated real-time image quality score is 60 points, below the threshold. The system automatically adjusts the resolution to 2560×1440 pixels (2K resolution), increases the brightness to 120 nits, and adjusts the contrast to 60%. After the adjustment, the image quality reaches 80 points, exceeding the threshold, resulting in a clear target video image.
[0082] The steps involved in extracting keyframes from the target video frame to obtain multiple keyframe images, extracting geographic information, spatial architectural information, and key graphic and textual information from the keyframe images, and generating a spatial semantic information set are as follows:
[0083] Perform frame similarity recognition on the target video frame to obtain the similarity value between each frame;
[0084] Frame recognition is performed on the image frames based on similarity values to determine the key frame positions. Key frames are then extracted based on the key frame positions to obtain multiple key frame images.
[0085] The objects in the keyframes are identified to obtain geographic information, spatial architectural information, and key graphic and textual information.
[0086] Geographic information, spatial architectural information, and key graphic and textual information are classified, integrated, and stored together to generate a spatial semantic information set.
[0087] In practice, the system extracts keyframes from the target video footage. The video contains 200 frames. The system calculates the similarity score between each frame (range 0.1~0.9), identifying frames with a similarity score below 0.3 as keyframes, specifically frames 30, 90, and 160. After extracting the keyframes, the system identifies objects in the footage: In frame 30, the park's iconic pavilion is identified, and spatial architectural information (e.g., octagonal roof), geographic information (e.g., lakeside coordinates), and key textual information (e.g., "No Swimming" sign) are extracted. In frame 90, a small bridge is identified, and spatial architectural information (stone arch structure), geographic information (bridge location), and key textual information ("Caution: Falling into Water" warning) are extracted. In frame 160, a bench is identified, and geographic information (bench coordinates) and key textual information ("Rest Area" sign) are extracted. Finally, the spatial architectural information, geographic information, and key textual information are integrated to generate a spatial semantic information set, providing detailed static information for location services.
[0088] The steps of extracting dynamic scene content from the target video frame, obtaining the dynamic scene subject and its actions based on the dynamic scene content, and obtaining dynamic semantic information based on the dynamic scene subject and its actions are as follows:
[0089] Based on the target video frame, dynamic recognition is performed on the target video frame to obtain the dynamic scene content in the target video frame;
[0090] Dynamic activity recognition is performed on dynamic scene content to obtain multiple dynamic scene subjects in the dynamic scene content;
[0091] Continuously monitor the main subject in a dynamic scene and capture the actions of the main subject during its movement.
[0092] Static feature recognition is performed on the main body in a dynamic scene to obtain the static features of the main body. Based on the static features of the main body and the actions of the main body in the scene, dynamic feature recognition is performed to obtain dynamic semantic information.
[0093] In application, the system extracts dynamic scene content from the target video footage. It identifies two main subjects in the dynamic scene: a runner wearing a yellow jacket (Subject A) and a woman pushing a stroller (Subject B). Subject A's movements are continuously monitored: stride length approximately 1 meter, speed approximately 8 km / h; Subject B's movements: slowly pushing the stroller, with stable hand gestures. Static features of Subject A are identified: height approximately 170 cm, slender build; Subject B: height approximately 160 cm, wearing a hat. Combining dynamic behaviors, Subject A is characterized as "running quickly," and Subject B is characterized as "pushing at a steady pace." The system integrates static attributes (height, clothing) and dynamic behaviors (action type, speed) to generate dynamic semantic information, providing crucial motion scene data for subsequent matching.
[0094] The steps for obtaining the target location range by performing a range query on a preset location map based on the spatial semantic information set are as follows:
[0095] Based on geographic information from spatial semantic information sets, a large-scale positioning pointer is generated; based on spatial architectural information, a medium-scale positioning pointer is generated; and based on key graphic and textual information, a small-scale positioning pointer is generated.
[0096] Based on the large-scale positioning pointer, the geographic information is matched on the preset positioning map to obtain the target's large-scale location;
[0097] Based on the mid-range pointer, spatial building information is matched within the target's large-range location to obtain the target's mid-range location;
[0098] Based on the small-range pointer, key graphic and textual information is matched at the target's range location to obtain the target's positioning range.
[0099] In practice, the system uses a set of spatial semantic information to gradually locate the target area: First, it generates a large-scale positioning pointer using geographic information (such as park coordinates), matching the park area (approximately 500 square meters). Second, it generates a medium-scale positioning pointer using spatial architectural information (such as the pavilion in the lake), matching the lakeside area (approximately 100 square meters) within the park area. Third, it generates a small-scale positioning pointer using key graphic information (such as "No Swimming" signs), matching the area near the signs in the lakeside area (approximately 20 square meters). Finally, the target positioning area is determined to be a 20-square-meter area southeast of the pavilion in the lake. This process achieves precise positioning by narrowing the scope layer by layer.
[0100] The steps for capturing comparison reference images from surveillance cameras within the target location range, comparing the spatial semantic information set and dynamic semantic information with the comparison reference images, and obtaining the comparison similarity are as follows:
[0101] Based on the target positioning range, extract data from all surveillance cameras within the target positioning range, and capture comparative reference images within the target positioning range based on the surveillance camera data;
[0102] Geographic information, spatial architectural information, and key graphic information from the spatial semantic information set are substituted into the comparison reference image for feature matching to obtain the first matching value;
[0103] Based on dynamic semantic information, video segments are searched in the comparison reference image to obtain the target comparison video segment, and the dynamic semantic information is matched with the frequency band of the target comparison video to obtain the second matching value;
[0104] The first matching value is combined with the second matching value to generate a comprehensive matching value. The similarity is then evaluated based on the comprehensive matching value to obtain the comparative similarity.
[0105] In practice, the system extracts data from three surveillance cameras (C1, C2, and C3) within the target location area (near the pavilion in the lake). After capturing comparison reference images from the cameras: First, the geographical information (lakeside coordinates) from the spatial semantic information set is substituted into the comparison images for matching, achieving a similarity of 85% (first matching value). Second, based on the dynamic semantic information (a runner in a yellow jacket), similar dynamic segments are found in the image from camera C1, achieving a similarity of 90% (second matching value). The first and second matching values are combined, calculating a comprehensive matching value of 87.5%, and the comparison similarity is evaluated as "high matching".
[0106] The steps of filtering multiple surveillance cameras based on similarity to obtain target cameras, extracting the camera location information of target cameras, and determining the alarm location based on the camera location information are as follows:
[0107] Multiple surveillance cameras were filtered based on their similarity to identify the target cameras closest to the crime scene.
[0108] Extract the location information, installation parameters, and camera footage from multiple target cameras;
[0109] The shooting direction of the target camera is obtained based on the installation parameters, and the relative distance and relative direction between the crime scene and the target camera are obtained based on the camera footage.
[0110] Determine the relative direction of the target between the crime scene and the target camera based on the relative direction and the shooting direction;
[0111] Determine the alarm location based on the target's relative direction and distance.
[0112] In practice, the system filters surveillance cameras based on similarity: C1 has a similarity of 87.5%, C2 75%, and C3 72%, with C1 selected as the target camera. The system extracts C1's location information (installed on a lamppost, 4 meters high), installation parameters (120-degree viewing angle), and real-time video feed. After analyzing the video, the relative distance between the crime scene and the camera is determined to be 25 meters, with a relative direction of northeast-east. Combined with the camera's shooting direction (due east), the target's relative direction is determined to be 15 degrees northeast-east. The final calculated alarm location coordinates are (latitude X, longitude Y), with an error of less than 3 meters.
[0113] After filtering multiple surveillance cameras based on similarity to determine the target camera closest to the crime scene, the process also includes:
[0114] Extract the surveillance footage from the target camera and perform a matching deviation assessment on the surveillance footage to obtain the deviation assessment value;
[0115] If the deviation assessment value is determined to be greater than the preset assessment threshold, then the camera resolution and camera position of multiple surveillance cameras are extracted.
[0116] The importance of each camera's location is assessed to obtain a location importance value for each target camera;
[0117] The image credibility value of each target camera is calculated by weighting the location importance value and the camera resolution. The target cameras are then further filtered based on the image credibility value.
[0118] In application, the system performs deviation assessment on the target camera (C1): extracting its real-time footage and performing a secondary match with the dynamic features of the alarm video; the deviation value is 15%. The preset deviation threshold is 10%, and since 15% > 10%, a secondary screening is initiated. The resolution (C1: 1080P, C2: 720P, C3: 720P) and location importance (C1: near the lake pavilion, importance value 90; C2: park entrance, value 70; C3: path, value 60) of all cameras are extracted. The image confidence level is calculated as follows: C1 = (90 × 0.6 + 1080 × 0.4) = 93.6, C2 = 74.8, C3 = 68.4. C1, with the highest confidence level, is selected, ultimately confirming the alarm location is correct.
[0119] After determining the alarm location based on the target's relative direction and distance, the process also includes:
[0120] The system acquires alarm voice information uploaded by citizens during the alarm reporting process, transcribes the alarm voice information, and extracts location-related sentences.
[0121] Semantic recognition is performed on location-related statements to obtain directional description information, location description information, and environmental description information, which are then combined to generate a speech location information set;
[0122] The voice location information set is weighted and then sent to the spatial semantic information set for information filtering and updating.
[0123] Based on the voice location information set, the alarm location is matched for location confirmation to obtain the confirmation matching value, and it is determined whether the confirmation matching value exceeds the preset confirmation matching threshold.
[0124] If it is determined that the confirmed matching value does not exceed the confirmed matching threshold, the monitoring camera and camera image are re-identified based on the voice location information set to obtain the identification result, and the alarm location is updated based on the identification result.
[0125] In practice, when a citizen dials 110 to report an emergency, in addition to video data, the system also obtains the voice information. The citizen says in the voice, "I was jogging on the east side of the pavilion in the middle of the lake and saw a child playing by the lake, with benches and several trees around." The system first transcribes this speech into text using a speech recognition API. Then, the system extracts location-related phrases from the text, including "east side of the pavilion in the middle of the lake," "lakeside," and "bench and several trees." Next, the system performs semantic recognition on these location-related phrases. After recognition, the system obtains the directional description "east side," the location description "near the pavilion in the middle of the lake," and the environmental description "lakeside, with benches and trees." The system aggregates this information to generate a voice location information set, containing three categories of descriptions: directional, location, and environment. Then, the system weights the voice location information set. Specifically, it assigns a weight of 0.3 to the directional description, 0.4 to the location description, and 0.3 to the environmental description (the total weight is 1). After weighting, the system sends this information set to the spatial semantic information set. The spatial semantic information set, previously generated from mobile phone camera data, includes coordinates, architectural features, and dynamic scenes within the park. The system filters and updates this set, retaining important information such as the coordinates of the pavilion in the center of the lake (latitude X1, longitude Y1) and lakeside environmental features, while deleting irrelevant data. Next, the system performs location confirmation matching against the previously determined alarm locations based on the voice location information set. The matching process includes calculating the similarity between the voice location information and the data in the spatial semantic information set. For example, the voice location information "near the pavilion in the center of the lake" matches the pavilion's coordinates (latitude X1, longitude Y1) in the spatial semantic information set, with a similarity score of 70%; "east side" matches the camera's viewing direction, with a similarity of 80%; and "lakeside, with benches and trees" matches environmental features, with a similarity of 75%. After comprehensive weighting, the confirmed matching value is (0.4 × 70% + 0.3 × 80% + 0.3 × 75%) = 74.5%. The system then determines whether this confirmed matching value exceeds the preset confirmation matching threshold of 80%. Because 74.5% is less than 80%, the system is deemed to have a match value that does not exceed the threshold. Since the match value is insufficient, the system re-identifies the surveillance cameras and their images based on the voice location information set. Specifically, the system uses the voice descriptions "east side of the pavilion," "lakeside," and "bench and trees" as keywords to search the real-time images of all available cameras (C1, C2, C3). For example, in the image from camera C1, the system identifies the pavilion but its direction is unclear; in the image from camera C2, the system finds the matching location of the bench and tree; in the image from camera C3, the system detects the running path and the eastward direction. After analysis, the identification result is: the best matching location is 10 degrees east of south of the pavilion, with coordinates updated to (latitude X2, longitude Y2), and an error of less than 5 meters.Finally, the system updates the alarm location based on this identification result, changing it from (latitude X1, longitude Y1) to (latitude X2, longitude Y2).
[0126] This invention provides an alarm location system based on 5G video multimodal large model analysis, using any of the alarm methods described above based on 5G video multimodal large model analysis. The system includes the following:
[0127] Video capture module 1: When a citizen uses a 5G mobile phone to report an emergency, it acquires the real-time video footage from the citizen's mobile phone camera and preprocesses the real-time video footage to obtain the target video footage.
[0128] Static feature module 2: used to extract keyframes from the target video frame, obtain multiple keyframe frames, extract geographic information, spatial architectural information and key graphic information from the keyframe frames, and generate a spatial semantic information set;
[0129] Dynamic Feature Module 3: Used to extract dynamic scene content from the target video frame, obtain the dynamic scene subject and scene subject action based on the dynamic scene content, and obtain dynamic semantic information based on the dynamic scene subject and scene subject action;
[0130] Initial positioning module 4: Used to perform range query on a preset positioning map based on the spatial semantic information set to obtain the target positioning range;
[0131] Image comparison module 5: Used to capture comparison reference images captured by surveillance cameras within the target positioning range, compare the spatial semantic information set and dynamic semantic information with the comparison reference images to obtain the comparison similarity;
[0132] Location determination module 6: It is used to filter multiple surveillance cameras based on comparison similarity to obtain the target camera, extract the camera location information of the target camera, and determine the alarm location based on the camera location information.
[0133] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. An alarm location method based on 5G video multimodal large model analysis, characterized in that, Includes the following steps: When a citizen uses a 5G mobile phone to report an emergency, the system acquires real-time video footage from the citizen's mobile phone camera and preprocesses the real-time video footage to obtain the target video footage. Keyframes are extracted from the target video frame to obtain multiple keyframe frames. Geographic information, spatial architectural information, and key graphic information are extracted from the keyframe frames to generate a spatial semantic information set. Extract dynamic scene content from the target video frame, obtain the dynamic scene subject and scene subject action based on the dynamic scene content, and obtain dynamic semantic information based on the dynamic scene subject and scene subject action; Based on the spatial semantic information set, a range query is performed on the preset positioning map to obtain the target positioning range; Capture the comparison reference images captured by the surveillance cameras within the target positioning range, compare the spatial semantic information set and the dynamic semantic information with the comparison reference images to obtain the comparison similarity; Multiple surveillance cameras are filtered based on the comparison similarity to obtain target cameras. The camera location information of the target cameras is extracted, and the alarm location is determined based on the camera location information.
2. The alarm location method based on 5G video multimodal large model analysis according to claim 1, characterized in that, The steps of acquiring real-time video footage from a citizen's mobile phone camera and preprocessing the real-time video footage to obtain the target video footage are as follows: When a citizen uses a 5G mobile phone to report an emergency, the system automatically captures the real-time video feed from the citizen's mobile phone camera. Extract the video resolution, video brightness, and video contrast of the real-time video frame, and obtain the real-time image quality of the real-time video frame based on the video resolution, video brightness, and video contrast. Determine whether the real-time image quality exceeds a preset quality threshold; if it is determined that the real-time image quality does not exceed the quality threshold. The video resolution, video brightness, and video contrast are then automatically adjusted to obtain a target video image whose real-time image quality exceeds the quality threshold.
3. The alarm location method based on 5G video multimodal large model analysis according to claim 2, characterized in that, The steps of extracting keyframes from the target video frame to obtain multiple keyframe frames, and extracting geographic information, spatial architectural information, and key graphic information from the keyframe frames to generate a spatial semantic information set are as follows: Frame similarity recognition is performed on the target video frame to obtain the similarity value between each frame; Frame recognition is performed on the image frames based on the similarity values to determine the key frame positions. Key frame extraction is then performed based on the key frame positions to obtain multiple key frame images. The objects in the keyframe images are identified to obtain geographic information, spatial architectural information, and key graphic information, respectively. The geographic information, spatial architectural information, and key graphic information are classified, integrated, and stored together to generate a spatial semantic information set.
4. The alarm location method based on 5G video multimodal large model analysis according to claim 3, characterized in that, The steps of extracting dynamic scene content from the target video frame, obtaining the dynamic scene subject and the scene subject's actions based on the dynamic scene content, and obtaining dynamic semantic information based on the dynamic scene subject and the scene subject's actions are as follows: Based on the target video frame, the target video frame is dynamically identified to obtain the dynamic scene content in the target video frame; Dynamic activity recognition is performed on the dynamic scene content to obtain multiple dynamic scene subjects in the dynamic scene content; The dynamic scene subject is continuously monitored, and the scene subject's actions during the movement of the dynamic scene subject are captured; Static feature recognition is performed on the subject of the dynamic scene to obtain the subject's static features. Based on the subject's static features and the subject's actions in the scene, dynamic feature recognition is performed to obtain dynamic semantic information.
5. The alarm location method based on 5G video multimodal large model analysis according to claim 4, characterized in that, The steps for obtaining the target location range by performing a range query on a preset location map based on the spatial semantic information set are as follows: Based on the geographic information in the spatial semantic information set, a large-scale positioning pointer is generated; based on the spatial building information, a medium-scale positioning pointer is generated; and based on the key graphic information, a small-scale positioning pointer is generated. Based on the large-scale positioning pointer, the geographic information is matched on the preset positioning map to obtain the target's large-scale location; Based on the mid-range pointer, the spatial building information is matched at the target's large-range location to obtain the target's mid-range location; Based on the small-range pointer, the key graphic information is matched at the range position in the target to obtain the target positioning range.
6. The alarm location method based on 5G video multimodal large model analysis according to claim 5, characterized in that, The steps of capturing comparison reference images from surveillance cameras within the target positioning range, comparing the spatial semantic information set and the dynamic semantic information with the comparison reference images to obtain the comparison similarity are as follows: Based on the target positioning range, extract data from all surveillance cameras within the target positioning range, and capture a comparative reference image within the target positioning range based on the surveillance camera data; Substitute the geographic information, the spatial architectural information, and the key graphic information from the spatial semantic information set into the comparison reference image for feature matching to obtain a first matching value; Based on the dynamic semantic information, video segments are searched in the comparison reference frame to obtain the target comparison video segment, and the dynamic semantic information is matched with the frequency band of the target comparison video to obtain a second matching value; The first matching value and the second matching value are combined to generate a comprehensive matching value. The similarity is then evaluated based on the comprehensive matching value to obtain the comparative similarity.
7. The alarm location method based on 5G video multimodal large model analysis according to claim 6, characterized in that, The steps of filtering multiple surveillance cameras based on the comparison similarity to obtain target cameras, extracting the camera location information of the target cameras, and determining the alarm location based on the camera location information are as follows: Based on the comparison similarity, multiple surveillance cameras are filtered to obtain multiple target cameras that are closest to the crime scene; Extract the location information, installation parameters, and camera images of multiple target cameras; The shooting direction of the target camera is obtained based on the installation parameters, and the relative distance and relative direction between the crime scene and the target camera are obtained based on the camera image. The relative direction between the crime scene and the target camera is determined based on the relative direction and the shooting direction. The alarm location is determined based on the relative direction and relative distance of the target.
8. The alarm location method based on 5G video multimodal large model analysis according to claim 7, characterized in that, After filtering multiple surveillance cameras based on the comparison similarity to obtain the target camera closest to the crime scene, the method further includes: Extract the monitoring footage from the target camera and perform a matching deviation assessment on the monitoring footage to obtain a deviation assessment value; If the deviation evaluation value is determined to be greater than the preset evaluation threshold, then the camera resolution and camera position of multiple surveillance cameras are extracted. The importance of the camera positions is evaluated to obtain a positional importance value for each target camera; A weighted average is calculated based on the location importance value and the camera resolution to obtain the image credibility value of each target camera. The target cameras are then further filtered based on the image credibility value.
9. The alarm location method based on 5G video multimodal large model analysis according to claim 7, characterized in that, After determining the alarm location based on the target's relative direction and relative distance, the method further includes: The system acquires alarm voice information uploaded by citizens during the alarm process, transcribes the alarm voice information, and extracts location-related sentences. Semantic recognition is performed on the location-related statements to obtain directional description information, location description information, and environmental description information, which are then combined to generate a voice location information set; The weighted set of voice location information is sent to the set of spatial semantic information, and the set of spatial semantic information is filtered and updated. Based on the voice location information set, the alarm location is matched for location confirmation to obtain a confirmation matching value, and it is determined whether the confirmation matching value exceeds a preset confirmation matching threshold. If it is determined that the confirmed matching value does not exceed the confirmed matching threshold, then the monitoring camera and the camera image are re-identified according to the voice location information set to obtain the identification result, and the alarm location is updated according to the identification result.
10. An alarm location system based on 5G video multimodal large model analysis, wherein the system uses the alarm location method based on 5G video multimodal large model analysis as described in any one of claims 1-9, characterized in that, The system includes: Video capture module: When a citizen uses a 5G mobile phone to report an emergency, it acquires the real-time video footage from the citizen's mobile phone camera and preprocesses the real-time video footage to obtain the target video footage. Static feature module: used to extract keyframes from the target video frame to obtain multiple keyframe frames, extract geographical information, spatial building information and key graphic information from the keyframe frames, and generate a spatial semantic information set; Dynamic feature module: used to extract dynamic scene content from the target video frame, obtain dynamic scene subject and scene subject action based on the dynamic scene content, and obtain dynamic semantic information based on the dynamic scene subject and scene subject action; Initial positioning module: used to perform range query on a preset positioning map based on the spatial semantic information set to obtain the target positioning range; Image comparison module: used to capture comparison reference images captured by surveillance cameras within the target positioning range, compare the spatial semantic information set and the dynamic semantic information with the comparison reference images to obtain the comparison similarity; Location determination module: used to filter multiple surveillance cameras according to the comparison similarity to obtain the target camera, extract the camera location information of the target camera, and determine the alarm location based on the camera location information.
Citation Information
Patent Citations
Emergency help-seeking method and system based on recognition and positioning of photos of mobile phone
CN102360433A
Online car-hailing passenger position rapid positioning method based on voice interaction and visual perspective
CN119399845A