A multi-modal multi-agent law enforcement assistance method for a railway police chest-worn law enforcement recorder

CN122802645APending Publication Date: 2026-09-22GUANGZHOU INST OF RAILWAY TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610945394.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,在上述智能功能叠加部署的过程中,现有设备普遍面临一个核心问题:各智能功能模块独立运行、互不通信,缺乏统一的协同机制来管理有限的计算、存储、网络及提示资源

Benefits of technology

[0053]通过将视频、音频、姿态、位置、设备状态和铁路业务信息统一纳入现场态势判断框架,使执法记录仪能够主动感知铁路执法现场的复杂状态,有效克服了单一画面或声音判断容易产生误报漏报的缺陷;通过风险评分与证据价值评估的联动机制,在冲突升级、民警摔倒或设备被抢夺等突发场景下自动锁定关键片段并生成证据标签,解决了民警因专注处置无法及时手动标记导致的证据遗漏问题;通过四个智能体的统一动作管理与资源冲突仲裁机制,使摄像头、存储、网络和提示通道等有限资源在民警安全、证据质量和设备续航之间得到有序协调,避免了功能模块独立运行造成的重复提示和资源抢占;通过将智能体触发原因、仲裁过程和民警确认结果写入证据元数据并与视音频文件共同封存,使事后审查不仅能够还原现场情况,还能够解释设备为何锁定某一片段、为何触发某种提示以及为何采用某种执行策略,从而显著提升了执法证据的完整性、可解释性和可追溯性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802645A_ABST
    Figure CN122802645A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal, multi-agent law enforcement assistance method for railway police chest-worn body cameras. The method collects on-site video, audio, posture, location, equipment status, and railway business information; generates on-site situational information through time alignment and scene normalization; extracts posture features and combines them with multimodal anomaly features for weighted fusion to assess risk levels; automatically locks key audio and video caches based on risk and evidentiary value and generates evidence markers; generates adaptive prompts based on risk and equipment status and schedules recording, storage, and uploading resources; four agents handle resource conflicts through unified arbitration and support police confirmation; finally, the risk assessment, protective actions, arbitration process, and confirmation results are written into evidence metadata and sealed together with key evidence files. This invention transforms the body camera from a passive recording device into an intelligent law enforcement terminal that actively senses, collaboratively protects, and is traceable, improving the integrity, interpretability, and traceability of evidence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of law enforcement recorder technology, specifically to a multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders. Background Technology

[0002] When railway police officers are on duty in stations, platforms, waiting rooms, security checkpoints, train carriages, and operating sections, the body-worn cameras they wear play a crucial role in capturing on-site audio and video and preserving evidence. With the continuous development of intelligent sensing technology, body-worn cameras are gradually integrating various intelligent analysis functions, such as video anomaly recognition, audio keyword detection, posture perception, key evidence marking, data upload scheduling, and on-site prompts, aiming to improve the timeliness and intelligence of law enforcement evidence collection. However, in the process of deploying these intelligent functions, existing equipment generally faces a core problem: the various intelligent function modules operate independently and do not communicate with each other, lacking a unified collaborative mechanism to manage limited computing, storage, network, and prompting resources.

[0003] Specifically, the video recognition module may request continuous high-bitrate recording, the posture perception module may trigger vibration alerts, the evidence marking module may require locking the cache and initiating uploads, and the prompting module may simultaneously issue voice warnings—these actions may simultaneously make conflicting resource requests to the camera, storage space, network channels, or prompting channels. However, existing equipment lacks a unified arbitration mechanism to determine whether to prioritize police safety, the quality of key evidence, device battery life, or low-interference prompts. Given the objective constraints of small size, limited battery power, and limited computing power of chest-worn devices, if resource conflicts cannot be coordinated in an orderly manner, it can easily lead to consequences such as repeated prompts interfering with police enforcement, key evidence being missed or covered due to resource contention, and inconsistent system judgments. Furthermore, traditional evidence documents only store audio and video, time, and device serial numbers. Information such as the analysis process generated by intelligent functions, resource conflict handling records, police confirmation or intervention actions, and the final execution strategy are not included in the evidence system. This makes it impossible to explain why the device locked a certain key segment, triggered a specific prompt, or adopted a certain saving or uploading strategy during post-event review, turning the intelligent analysis process into a "black box" and failing to support the requirements of the integrity and traceability of law enforcement evidence. Therefore, how to construct an enforcement assistance method that can unify and coordinate multiple intelligent functions, orderly arbitrate resource conflicts, and fully record, analyze, and enforce the entire process has become an urgent technical problem to be solved in this field. Summary of the Invention

[0004] To address the technical problems mentioned in the background section, this invention provides a multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders.

[0005] The above-mentioned objective of this application is achieved through the following technical solution:

[0006] A multimodal, multi-agent law enforcement assistance method for railway police chest-worn body camera devices includes the following steps:

[0007] S10: Collect multimodal raw perception information from the railway law enforcement site, perform time alignment and scene normalization processing on the multimodal raw perception information, and generate on-site situation information;

[0008] S20: Extract attitude feature information based on the on-site situation information, and perform a comprehensive risk assessment by combining the on-site situation information and the attitude feature information to generate risk score information and risk level information;

[0009] S30: Based on the risk score information and the risk level information, the evidence value is assessed. When the assessment result meets the preset conditions, a key evidence protection action is triggered to generate key evidence marking information and lock the corresponding audio and video cache information.

[0010] S40: Generate handling assistance prompts based on the risk level information, the key evidence marking information, and the equipment status information, and select the corresponding prompt method according to the current on-site environment status;

[0011] S50: Based on the risk level information, the key evidence marking information, and the equipment resource status information, make resource scheduling decisions and generate resource allocation instruction information;

[0012] S60: The agent action request information generated in steps S10 to S50 is uniformly organized and conflict arbitration is performed to generate the final execution action information;

[0013] S70: Write the risk scoring information, the key evidence marking information, the resource allocation instruction information, the conflict arbitration process information, and the police confirmation result information into the evidence metadata information to form a traceable law enforcement evidence closed loop.

[0014] In a preferred embodiment, this application can be further configured such that step S10 specifically includes:

[0015] S101: Collect on-site video information, on-site audio information, equipment posture information, location information, equipment operating status information, and railway business information through a chest-worn law enforcement recorder to generate the multimodal raw perception information;

[0016] S102: Time synchronization and alignment of various types of information in the original multimodal sensing information, and filtering of abnormal data to generate synchronized multimodal sensing information;

[0017] S103: Perform scene normalization processing on the synchronized multimodal perception information to generate the on-site situation information.

[0018] In a preferred embodiment, this application can be further configured such that step S20 specifically includes:

[0019] S201: Extract acceleration feature information and angular velocity feature information from the equipment attitude information in the on-site situation information, and generate the attitude feature information;

[0020] S202: The video anomaly feature information, audio anomaly feature information, location sensitivity information, equipment anomaly status information and attitude feature information in the on-site situation information are weighted and fused to generate the risk score information;

[0021] S203: Compare the risk score information with the preset risk level threshold information to generate the risk level information.

[0022] In a preferred embodiment, this application can be further configured such that step S30 specifically includes:

[0023] S301: Based on the aforementioned risk level information, event type information, audio and video integrity information, and railway scene location information, a comprehensive score of evidence value is generated to produce evidence value score information;

[0024] S302: When the evidence value scoring information reaches the preset evidence protection trigger threshold, the audio and video cache information within a preset time period before and after the event is automatically locked, and key evidence marking information containing event time, location and event type information is generated;

[0025] S303: Perform anti-overwriting protection and integrity verification on the audio and video cache information corresponding to the key evidence marking information, and generate evidence protection status information.

[0026] In a preferred embodiment, this application can be further configured such that step S40 specifically includes:

[0027] S401: Based on the risk level information, the key evidence marking information, the police officer's posture and status information, and the equipment status information, generate auxiliary prompt information for handling that includes the prompt type and prompt intensity;

[0028] S402: Obtain current ambient noise intensity information and ambient lighting conditions information, and select a corresponding prompting method based on the ambient noise intensity information and ambient lighting conditions information. The prompting method includes vibration prompting, indicator light prompting, text prompting, or voice prompting.

[0029] In a preferred embodiment, this application can be further configured such that step S50 specifically includes:

[0030] S501: Obtain current device resource status information, which includes power information, storage space information, network status information, and computing load information;

[0031] S502: Determine task priority information based on the risk level information and the key evidence marking information;

[0032] S503: Generate resource allocation instruction information based on the task priority information and the device resource status information. The resource allocation instruction information includes recording quality adjustment instruction, storage space allocation instruction, network upload scheduling instruction, and prompt channel allocation instruction.

[0033] In a preferred embodiment, this application can be further configured such that step S60 specifically includes:

[0034] S601: Obtain action request information generated by the situational awareness agent, evidence protection agent, handling assistance agent, and resource scheduling agent respectively. The action request information includes action source information, target resource information, event level information, and priority information.

[0035] S602: Based on the action request information, query the resource control table information to obtain the current resource occupancy status information and occupancy source information;

[0036] S603: When multiple action request messages conflict for the same resource, the final execution action information is generated according to the preset arbitration rules. The preset arbitration rules include the police safety priority rule, the key evidence protection priority rule, the high-risk event priority rule, and the low-interference prompt priority rule.

[0037] S604: When the final execution action information involves law enforcement judgment or event escalation, generate police confirmation request information and update the final execution action information based on the police confirmation feedback information.

[0038] In a preferred embodiment, this application can be further configured such that step S70 specifically includes:

[0039] S701: The risk scoring information, the key evidence marking information, the resource allocation instruction information, the conflict arbitration process information, and the police confirmation result information are summarized to generate evidence metadata information;

[0040] S702: Associate the evidence metadata information with the corresponding key evidence files and perform integrity verification on the evidence metadata information and the key evidence files to generate evidence integrity verification information;

[0041] S703: Seal the evidence metadata information together with the key evidence documents to form a traceable closed loop of law enforcement evidence.

[0042] In a preferred embodiment, this application can be further configured such that the railway business information includes station information, platform information, train number information, carriage information, operating section information, and duty task type information.

[0043] The second objective of this invention is achieved through the following technical solution:

[0044] A multimodal, multi-agent law enforcement assistance system for railway police chest-worn body camera systems includes:

[0045] The multimodal data acquisition module is used to collect multimodal raw perception information from railway law enforcement sites and perform time alignment and scene normalization processing to generate on-site situation information;

[0046] The railway law enforcement situation awareness intelligent agent module is used to conduct a comprehensive risk assessment based on the on-site situation information and attitude feature information, and generate risk score information and risk level information.

[0047] The key evidence identification and protection intelligent agent module is used to assess the value of evidence based on the risk score information and the risk level information, trigger key evidence protection actions, and generate key evidence marking information.

[0048] The law enforcement handling assistance intelligent agent module is used to generate handling assistance prompts based on the risk level information, the key evidence marking information, and the equipment status information.

[0049] The resource scheduling intelligent agent module is used to make resource scheduling decisions based on the risk level information, the key evidence marking information, and the equipment resource status information, and to generate resource allocation instruction information.

[0050] The conflict arbitration module is used to uniformly organize and arbitrate the action request information generated by each intelligent agent module, and generate the final execution action information.

[0051] The evidence metadata management module is used to write risk assessment information, key evidence marking information, resource allocation instruction information, conflict arbitration process information, and police confirmation result information into the evidence metadata information, forming a traceable closed loop of law enforcement evidence.

[0052] The beneficial effects of the multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders of this invention are as follows:

[0053] By integrating video, audio, posture, location, equipment status, and railway business information into a unified on-site situation assessment framework, the law enforcement recorder can proactively perceive the complex state of railway law enforcement scenes, effectively overcoming the shortcomings of relying solely on visual or auditory judgments that are prone to false alarms or omissions. Through a linkage mechanism between risk scoring and evidence value assessment, it automatically locks key segments and generates evidence tags in emergency scenarios such as escalating conflicts, police officers falling, or equipment being seized, solving the problem of evidence omissions caused by police officers being unable to manually mark evidence in time due to their focus on handling the situation. Through unified action management of four intelligent agents and a resource conflict arbitration mechanism, limited resources such as cameras, storage, networks, and prompting channels are coordinated in an orderly manner between police safety, evidence quality, and equipment battery life, avoiding duplicate prompts and resource contention caused by independent operation of functional modules. By writing the triggering reasons of intelligent agents, arbitration processes, and police confirmation results into evidence metadata and sealing them together with audio and video files, post-event review can not only reconstruct the scene but also explain why the equipment locked a certain segment, why a certain prompt was triggered, and why a certain execution strategy was adopted, thereby significantly improving the integrity, interpretability, and traceability of law enforcement evidence. Attached Figure Description

[0054] Figure 1 This is a flowchart of an embodiment of a multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders;

[0055] Figure 2 This is a flowchart illustrating step S10 in an embodiment of a multimodal, multi-agent law enforcement assistance method for a chest-worn law enforcement recorder for railway police, as described in this application.

[0056] Figure 3 This is a flowchart illustrating step S20 in an embodiment of a multimodal, multi-agent law enforcement assistance method for a chest-worn law enforcement recorder for railway police, as described in this application.

[0057] Figure 4 This is a flowchart illustrating step S30 in an embodiment of a multimodal, multi-agent law enforcement assistance method for a chest-worn law enforcement recorder for railway police, as described in this application.

[0058] Figure 5 This is a flowchart illustrating step S50 in an embodiment of a multimodal, multi-agent law enforcement assistance method for a chest-worn law enforcement recorder for railway police, as described in this application.

[0059] Figure 6 This is a flowchart illustrating step S60 in an embodiment of a multimodal, multi-agent law enforcement assistance method for a chest-worn law enforcement recorder for railway police, as described in this application.

[0060] Figure 7This is a flowchart illustrating step S70 in an embodiment of a multimodal, multi-agent law enforcement assistance method for a chest-worn law enforcement recorder for railway police, as described in this application. Detailed Implementation

[0061] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] As attached Figure 1-7 As shown, a multimodal, multi-agent law enforcement assistance method for railway police chest-worn body cameras includes the following steps:

[0063] S10: Collect multimodal raw perception information from the railway law enforcement site, perform time alignment and scene normalization processing on the multimodal raw perception information, and generate on-site situation information;

[0064] S20: Extract attitude feature information based on the on-site situation information, and perform a comprehensive risk assessment by combining the on-site situation information and the attitude feature information to generate risk score information and risk level information;

[0065] S30: Based on the risk score information and the risk level information, the evidence value is assessed. When the assessment result meets the preset conditions, a key evidence protection action is triggered to generate key evidence marking information and lock the corresponding audio and video cache information.

[0066] S40: Generate handling assistance prompts based on the risk level information, the key evidence marking information, and the equipment status information, and select the corresponding prompt method according to the current on-site environment status;

[0067] S50: Based on the risk level information, the key evidence marking information, and the equipment resource status information, make resource scheduling decisions and generate resource allocation instruction information;

[0068] S60: The agent action request information generated in steps S10 to S50 is uniformly organized and conflict arbitration is performed to generate the final execution action information;

[0069] S70: Write the risk scoring information, the key evidence marking information, the resource allocation instruction information, the conflict arbitration process information, and the police confirmation result information into the evidence metadata information to form a traceable law enforcement evidence closed loop.

[0070] In this embodiment, multimodal raw sensing information is a collection of unprocessed raw data collected by various sensors, including video, audio, posture, location, equipment status, and railway business information; time alignment is the operation of synchronizing sensing data from different sources according to the same time reference; scene normalization is the process of converting sensing data of different formats and dimensions into a unified expression form; on-site situation information is structured data that reflects the overall state of the current law enforcement scene after processing; posture feature information is feature data extracted from equipment motion sensors that reflects the physical state of police officers and the safety status of equipment; comprehensive risk assessment is the process of quantitatively calculating on-site safety risks based on multi-dimensional information; risk score information is a quantitative value reflecting the degree of on-site risk; risk level information is a risk level identifier divided according to risk score; evidence value assessment is the process of judging whether event fragments have preservation value; key evidence protection action is the operation of automatically locking and saving key audio and video fragments when conditions are met; key evidence The tagging information consists of labels used to mark locked evidence segments; the audio and video cache information consists of video and audio data temporarily stored by the law enforcement recorder that has not yet been finalized; the handling assistance prompt information consists of auxiliary decision-making prompts provided to police officers based on on-site risks; the device status information consists of data reflecting the current operating status of the law enforcement recorder, including battery level, storage, network, and recording status; the resource scheduling decision is the process of rationally allocating limited device resources; the resource allocation instruction information consists of control instructions instructing each functional module on how to allocate and use device resources; the intelligent agent action request information consists of operation request data issued by each intelligent module when performing tasks; the conflict arbitration is the process of adjudicating conflicts when multiple intelligent agents simultaneously request the same resource; the final execution action information consists of the operation instructions determined after arbitration; the evidence metadata information consists of descriptive data describing the attributes, formation process, and execution records of evidence files; and the law enforcement evidence closed loop is a complete and traceable link from on-site perception to risk identification, evidence protection, execution records, and post-event review.

[0071] Specifically, firstly, multimodal raw sensory information from railway law enforcement scenes is collected using body cameras. This data from different sensors undergoes time alignment and scene normalization, unifying previously scattered video, audio, posture, location, and equipment information into a single judgment framework to generate comprehensive situational information reflecting the scene. Then, posture feature information is extracted based on the situational information, and a comprehensive risk assessment is performed by combining the situational information and posture feature information to generate risk score and risk level information. This incorporates on-site video, audio, posture anomalies, location sensitivity, and railway business scenarios into the risk assessment, avoiding misjudgments or omissions caused by a single information source. Next, evidence value is assessed based on the risk score and risk level information. When the assessment results meet preset conditions, key evidence protection actions are automatically triggered, generating key evidence marker information and locking corresponding audio and video cache information, ensuring that key segments are automatically saved when police officers cannot manually operate the system. Simultaneously, based on risk level information and key evidence markers… Information and equipment status information generate auxiliary prompts for handling, and selects appropriate prompting methods based on the current on-site environment to ensure that the prompts are effective and do not interfere with law enforcement. Subsequently, resource scheduling decisions are made based on risk level information, key evidence marking information, and equipment resource status information, generating resource allocation instructions to prioritize critical tasks when power, storage, and computing power are limited. Next, the intelligent agent action request information generated in each step is uniformly organized and conflict arbitration is conducted to generate final execution action information, avoiding resource contention and decision inconsistencies caused by independent operation of various intelligent functions. Finally, risk scoring information, key evidence marking information, resource allocation instructions, conflict arbitration process information, and police confirmation results are written into the evidence metadata information, forming a traceable law enforcement evidence closed loop. This allows for post-event review not only to confirm what happened on-site but also to explain why the equipment locked a certain segment, why a certain prompt was triggered, and why a certain execution strategy was adopted, improving the completeness, interpretability, and traceability of law enforcement evidence.

[0072] In one embodiment, step S10 specifically includes:

[0073] S101: Collect on-site video information, on-site audio information, equipment posture information, location information, equipment operating status information, and railway business information through a chest-worn law enforcement recorder to generate the multimodal raw perception information;

[0074] S102: Time synchronization and alignment of various types of information in the original multimodal sensing information, and filtering of abnormal data to generate synchronized multimodal sensing information;

[0075] S103: Perform scene normalization processing on the synchronized multimodal perception information to generate the on-site situation information.

[0076] In this embodiment, the on-site video information is a continuous image sequence data collected by the law enforcement recorder camera; the on-site audio information is the environmental sound signal data collected by the law enforcement recorder microphone; the device attitude information is the angular velocity and acceleration data of the law enforcement recorder in three-dimensional space obtained by the inertial measurement unit; the location information is the spatial coordinate data of the law enforcement recorder obtained by the positioning module; the device operating status information is data reflecting the power, storage, network, recording status and sensor working status of the law enforcement recorder; the railway business information is business data related to the railway law enforcement scenario, including stations, platforms, train numbers, carriages, operating sections and duty task types, etc.; the synchronized multimodal perception information is a multi-source data set after time alignment and anomaly filtering; the scene normalization processing is to convert perception data from different sources and in different formats into a unified expression structure and dimensional standard.

[0077] Specifically, firstly, a chest-worn law enforcement recorder collects on-site video, audio, equipment posture, location, equipment operating status, and railway business information to generate multimodal raw perception information. This information includes video footage reflecting crowd gathering and physical movements, audio reflecting arguments and impact sounds, posture reflecting police officers running or falling, location and railway business information determining the specific law enforcement scenario, and equipment status reflecting the recorder's availability. These six types of information collectively form the foundation for complete on-site perception. Next, the various types of information in the multimodal raw perception information are synchronized in time, ensuring that video frames, audio sampling points, posture data, and location data are on the same time reference. Abnormal data is filtered to remove sensor noise and erroneous readings, generating synchronized multimodal perception information that provides a reliable data foundation for subsequent judgments. Finally, the synchronized multimodal perception information undergoes scene normalization processing, unifying data of different dimensions and formats into structured data that the system can process uniformly, generating on-site situational information. This step unifies the previously scattered sensor data, equipment data, and railway business data into a single judgment framework, enabling subsequent intelligent agents to conduct risk analysis based on complete, consistent, and multi-dimensional on-site situation data. This provides a reliable data foundation for accurately distinguishing between ordinary congestion, escalating conflicts, equipment malfunctions, and police safety risks.

[0078] In one embodiment, step S20 specifically includes:

[0079] S201: Extract acceleration feature information and angular velocity feature information from the equipment attitude information in the on-site situation information, and generate the attitude feature information;

[0080] S202: The video anomaly feature information, audio anomaly feature information, location sensitivity information, equipment anomaly status information and attitude feature information in the on-site situation information are weighted and fused to generate the risk score information;

[0081] S203: Compare the risk score information with the preset risk level threshold information to generate the risk level information.

[0082] In this embodiment, the acceleration feature information is the three-axis acceleration change data extracted from the equipment attitude information, used to determine whether the police officers are running, falling, pushing, or violently shaking; the angular velocity feature information is the three-axis rotational angular velocity data extracted from the equipment attitude information, used to determine whether the equipment is rapidly rotating or being robbed; the video anomaly feature information is the abnormal indicator data extracted from the on-site video information reflecting screen obstruction, violent shaking, crowd gathering, or conflict actions; the audio anomaly feature information is the acoustic feature data extracted from the on-site audio information reflecting arguments, cries for help, impact sounds, or abnormally high noise; the location sensitivity information is the risk weighting coefficient data determined according to the railway scene to which the event occurred; the equipment abnormal status information is the identification data reflecting equipment abnormalities such as camera obstruction, microphone failure, low battery, or insufficient storage; the weighted fusion calculation is the process of comprehensively calculating the abnormal features of different modalities according to their respective weights; and the preset risk level threshold information is the set of boundary values ​​pre-set by the system for classifying risk levels.

[0083] Specifically, firstly, acceleration and angular velocity features are extracted from equipment attitude information in the on-site situation information to generate attitude feature information. This enables the system to detect whether police officers are running, falling, pushing or shoving, equipment is shaking violently, or equipment is being snatched, thus retaining the ability to assist in judging emergencies even when the footage is unstable. Then, video anomaly features, audio anomaly features, location sensitivity information, equipment anomaly status information, and attitude feature information in the on-site situation information are weighted and fused to generate risk score information. This allows arguments in the carriage and crowding in the waiting hall to receive different risk scores due to different location sensitivity, thus accurately reflecting the impact of different railway scenarios on risk. Finally, the risk score information is compared with preset risk level threshold information to generate risk level information, enabling the system to classify events into levels such as ordinary, attention, important, or urgent based on the degree of risk. Through this step, the system no longer relies on a single video or audio recording to determine the situation on site. Instead, it uses multimodal weighted evaluation to enable the equipment to proactively identify law enforcement-significant events such as arguments in train carriages, pushing and shoving at platform edges, police officers falling, equipment obstruction, and abnormal collisions. This provides a unified and accurate risk entry point for subsequent evidence protection and handling.

[0084] In one embodiment, step S30 specifically includes:

[0085] S301: Based on the aforementioned risk level information, event type information, audio and video integrity information, and railway scene location information, a comprehensive score of evidence value is generated to produce evidence value score information;

[0086] S302: When the evidence value scoring information reaches the preset evidence protection trigger threshold, the audio and video cache information within a preset time period before and after the event is automatically locked, and key evidence marking information containing event time, location and event type information is generated;

[0087] S303: Perform anti-overwriting protection and integrity verification on the audio and video cache information corresponding to the key evidence marking information, and generate evidence protection status information.

[0088] In this embodiment, the comprehensive evidence value score is a multi-dimensional quantitative calculation process to determine whether an event segment has preservation value; the event type information is classification data identifying the category to which the current event belongs, including conflict events, equipment malfunction events, police safety events, etc.; the audio and video integrity information is quality assessment data reflecting whether the current audio and video segment is complete and whether there is any image obstruction or sound loss; the railway scene location information is location label data identifying the specific scene where the event occurs, such as a station, platform, waiting room, security checkpoint, train carriage, or operating section; the evidence value score information is a quantitative value reflecting the preservation value of the event segment evidence; the preset evidence protection trigger threshold is a scoring threshold preset by the system to trigger evidence protection actions; the audio and video cache information is temporary video and audio data that has not yet been finally saved in the law enforcement recorder's cyclic cache; the event time is the time identifier of the event occurrence; and the key evidence marking information is label data containing attributes such as event type, time, and location that annotates the locked evidence segments.

[0089] Specifically, the system first performs a comprehensive evaluation of evidence value based on risk level information, event type information, audio and video integrity information, and railway scene location information. This generates an evidence value score, enabling the system to comprehensively assess whether events such as platform edge conflicts, escalating arguments in carriages, police officer falls, or equipment being stolen have preservation value. It also assesses the integrity of images with severe obstruction or missing sound. When the evidence value score reaches a preset evidence protection trigger threshold, the system automatically locks the audio and video cache information within a preset time period before and after the event, generating key evidence marker information containing the event time, location, and event type. This ensures that key segments are automatically saved when police officers are focused on handling the situation and cannot manually mark them, solving the problem of untimely manual marking in emergency law enforcement scenarios. Then, the system performs anti-overwriting protection and integrity verification on the audio and video cache information corresponding to the key evidence marker information, generating evidence protection status information to ensure that key evidence segments are not overwritten by ordinary loop recordings and that their integrity can be verified during subsequent export, upload, and review. Through this process, risk assessment, evidence locking, tag generation, anti-overwriting protection, and integrity verification form a complete linkage mechanism, enabling evidence protection to move beyond the alarm level and achieve a closed loop from identification to locking.

[0090] In one embodiment, step S40 specifically includes:

[0091] S401: Based on the risk level information, the key evidence marking information, the police officer's posture and status information, and the equipment status information, generate auxiliary prompt information for handling that includes the prompt type and prompt intensity;

[0092] S402: Obtain current ambient noise intensity information and ambient lighting conditions information, and select a corresponding prompting method based on the ambient noise intensity information and ambient lighting conditions information. The prompting method includes vibration prompting, indicator light prompting, text prompting, or voice prompting.

[0093] In this embodiment, the prompt type is the identifier data that identifies the category of auxiliary prompt methods, including vibration prompts, indicator light prompts, text prompts, and voice prompts; the prompt intensity is the level data that identifies the urgency of the prompt, used to distinguish between general attention, important reminders, and emergency warnings; the police officer's posture status information is data reflecting the police officer's current physical state, including whether they are running, stationary, or have fallen; the on-site environmental noise intensity information is quantitative data reflecting the current environmental sound intensity, used to determine whether the voice prompt can be effectively received; and the on-site lighting conditions information is data reflecting the current environmental brightness, used to determine whether the visual prompt can be clearly perceived.

[0094] Specifically, firstly, based on risk level information, key evidence marking information, police officer posture and status information, and equipment status information, the system generates auxiliary prompts that include the type and intensity of the prompts. This allows the system to alert police officers to risks, confirm events, protect evidence, or request support at critical junctures. Simultaneously, it determines whether the prompts are likely to be perceived based on the officer's current state. Next, it acquires information on the current ambient noise level and lighting conditions. Based on these conditions, it selects the appropriate prompting method. In environments with high noise levels, such as station announcements or train operation noise, it reduces voice prompts to avoid ineffective prompts or interference with law enforcement communication. In well-lit conditions, it prioritizes indicator lights or text prompts. When police safety is threatened or the conflict escalates significantly, it triggers a higher-level alert and requests police confirmation. Through this process, the prompting strategy is matched to the ambient environment. The system does not simply trigger an alarm immediately after identifying an anomaly; instead, it selects an appropriate prompting method based on the risk level, ambient noise, police officer status, and equipment conditions, thereby significantly reducing the impact of false alarms and highly interfering prompts on on-site law enforcement.

[0095] In one embodiment, step S50 specifically includes:

[0096] S501: Obtain current device resource status information, which includes power information, storage space information, network status information, and computing load information;

[0097] S502: Determine task priority information based on the risk level information and the key evidence marking information;

[0098] S503: Generate resource allocation instruction information based on the task priority information and the device resource status information. The resource allocation instruction information includes recording quality adjustment instruction, storage space allocation instruction, network upload scheduling instruction, and prompt channel allocation instruction.

[0099] In this embodiment, the device resource status information is a data set reflecting the total amount of available resources of the law enforcement recorder, including power information, storage space information, network status information, and computing load information; power information reflects the device's current remaining power; storage space information reflects the device's current remaining storage capacity; network status information reflects the current wireless network connection quality and available bandwidth; computing load information reflects the device's current processor and memory usage; task priority information is an urgency indicator assigned to each task based on the event risk level and evidentiary value; recording quality adjustment instructions are control commands to adjust video recording resolution, frame rate, or bitrate; storage space allocation instructions are control commands indicating how to allocate limited storage resources to different data types; network upload scheduling instructions are control commands indicating when to upload, what type of data to upload, and what compression rate to use; and prompt channel allocation instructions are control commands indicating what prompt method to use and when to use the prompt channel.

[0100] Specifically, the system first acquires current device resource status information, including battery level, storage space, network status, and computing load, enabling it to understand the total available resources in real time. Then, based on risk level information and key evidence marking information, task priority is determined, prioritizing evidence protection tasks for high-risk events and lower-risk or routine tasks. Resource allocation instructions are generated based on task priority and device resource status information, including video quality adjustment instructions, storage space allocation instructions, network upload scheduling instructions, and alert channel allocation instructions. In emergencies, priority is given to preserving original evidence, locking key segments, and providing safety alerts to police officers. When network conditions are good, complete evidence files are uploaded; when network conditions are average, key segments or keyframes are uploaded first; when the network is poor or down, complete evidence is stored locally and uploaded later when the network is restored; when battery or storage space is insufficient, identified key evidence is protected and the use of resources by ordinary video recordings is limited. Through this process, the device no longer allocates resources evenly but establishes a key task-priority operating mechanism, significantly reducing the risk of evidence loss and core function interruption.

[0101] In one embodiment, step S60 specifically includes:

[0102] S601: Obtain action request information generated by the situational awareness agent, evidence protection agent, handling assistance agent, and resource scheduling agent respectively. The action request information includes action source information, target resource information, event level information, and priority information.

[0103] S602: Based on the action request information, query the resource control table information to obtain the current resource occupancy status information and occupancy source information;

[0104] S603: When multiple action request messages conflict for the same resource, the final execution action information is generated according to the preset arbitration rules. The preset arbitration rules include the police safety priority rule, the key evidence protection priority rule, the high-risk event priority rule, and the low-interference prompt priority rule.

[0105] S604: When the final execution action information involves law enforcement judgment or event escalation, generate police confirmation request information and update the final execution action information based on the police confirmation feedback information.

[0106] In this embodiment, the situational awareness agent is a functional module responsible for comprehensively analyzing the multimodal situational awareness and generating risk assessments; the evidence protection agent is a functional module responsible for judging the value of evidence and triggering key evidence protection actions; the handling assistance agent is a functional module responsible for generating handling assistance prompts and selecting the prompting method according to the environment; the resource scheduling agent is a functional module responsible for making resource allocation decisions based on event level and device status; the action source information is identification data that identifies which agent issued the action request; the target resource information is data that identifies whether the action request targets specific resources such as cameras, microphones, storage space, network channels, or prompting channels; the event level information is data that identifies the risk level of the event associated with the action request; the priority information is data that identifies the urgency of processing the action request in the current system; and the resource control table information records which agent currently controls each device resource. The system includes: a mapping table of resource occupancy and its priority; current resource occupancy status information indicating whether each device resource is currently occupied and by which task; occupancy source information identifying the agent or module currently occupying a resource; preset arbitration rules a set of pre-set arbitration rules for resolving resource conflicts; police safety priority rules prioritizing actions when police personal safety is involved; key evidence protection priority rules prioritizing evidence locking and preservation actions when key evidence is involved; high-risk event priority rules prioritizing actions related to high-risk events; low-interference prompt priority rules prioritizing prompts that minimize interference with on-site law enforcement when prompting police officers; and police confirmation request information sent to police officers when law enforcement judgments or event escalation are involved.

[0107] Specifically, the system first acquires action request information generated by the situational awareness agent, evidence protection agent, handling assistance agent, and resource scheduling agent. Each action request includes attributes such as action source, target resource, event level, and priority, ensuring that the outputs of each agent are not executed independently but enter a unified action management process. This avoids the problems of duplicate prompts and inconsistent decisions caused by the independent operation of each functional module. Based on the action request information, the system queries the resource control table to obtain the current resource occupancy status and occupancy source information, enabling the system to accurately grasp the current usage status and control source of resources such as cameras, microphones, storage space, network channels, and prompt channels. When multiple action requests conflict for the same resource, the system generates the final execution action information according to preset arbitration rules. These arbitration rules include priority for police safety, priority for key evidence protection, priority for high-risk events, and priority for low-interference prompts. The system does not simply overwrite the previous instruction with a subsequent instruction but makes an orderly decision. When the final execution action information involves law enforcement judgment or event escalation, the system generates a police confirmation request and updates the final execution action information based on the police confirmation feedback. In emergency situations, the system can execute safety or evidence protection actions first and record them afterward. This step enables orderly collaboration among multiple intelligent agents, resolving resource conflicts under limited resource conditions in chest-worn devices and allowing the system to make reasonable decisions based on the urgency of events and device status.

[0108] In one embodiment, step S70 specifically includes:

[0109] S701: The risk scoring information, the key evidence marking information, the resource allocation instruction information, the conflict arbitration process information, and the police confirmation result information are summarized to generate evidence metadata information;

[0110] S702: Associate the evidence metadata information with the corresponding key evidence files and perform integrity verification on the evidence metadata information and the key evidence files to generate evidence integrity verification information;

[0111] S703: Seal the evidence metadata information together with the key evidence documents to form a traceable closed loop of law enforcement evidence.

[0112] In this embodiment, the evidence metadata information is a comprehensive descriptive data describing all intelligent analysis and execution records throughout the entire process from identification to sealing of key evidence files; the key evidence files are audio and video data files that are locked by the system and marked as requiring protection; the integrity verification is the process of verifying whether the evidence files and their metadata have been tampered with using a hash algorithm; the evidence integrity verification information is the verification result data that identifies whether the evidence files and metadata are complete; sealing is the operation of marking the evidence files and metadata together as read-only and preventing modification or deletion.

[0113] Specifically, the process begins by summarizing risk scoring information, key evidence marking information, resource allocation instruction information, conflict arbitration process information, and police confirmation results to generate evidence metadata. This metadata records not only traditional time, location, device number, and file information, but also process information such as why the system was triggered during the event, which agent triggered it, what resource conflict occurred, what actions were ultimately performed, and whether the police confirmed or covered the event. Next, the evidence metadata is associated with and stored with the corresponding key evidence files. The integrity of the metadata and key evidence files is then verified, generating evidence integrity verification information to ensure that the key evidence files and intelligent analysis records can be confirmed as unaltered during subsequent review. Finally, the evidence metadata and key evidence files are sealed together, forming a traceable closed loop of law enforcement evidence. Through this process, traditional law enforcement videos can only prove what happened at the scene, while this invention can also explain why the device locked the segment, why a prompt was triggered, and why a certain upload or save strategy was adopted. This forms a complete closed loop from scene perception, risk identification, evidence protection, resource coordination, police confirmation to post-event review, significantly improving the integrity, interpretability, and traceability of law enforcement evidence.

[0114] In one embodiment, the railway business information includes station information, platform information, train number information, carriage information, operating section information, and duty task type information.

[0115] In this embodiment, railway business information is a set of business data related to railway law enforcement scenarios; station information is data identifying the station where the event occurred; platform information is data identifying the platform where the event occurred; train number information is data identifying the train number involved in the event; carriage information is data identifying the carriage number involved in the event; operating section information is data identifying the section of the line where the event occurred; and duty task type information is data identifying the type of duty task currently performed by the police officer, including patrol duty, guard duty, security check duty, and train-accompanying duty, etc.

[0116] Specifically, railway business information includes station information, platform information, train number information, carriage information, operating section information, and duty task type information. This allows the system to make comprehensive judgments not only based on physical perception data such as video, audio, and posture when assessing on-site risks, but also by combining this with the specific railway scenario in which the incident occurs. For example, when people gather or push, the risk level receives different sensitivity weights depending on whether it occurs in a waiting hall, security checkpoint, platform edge, or carriage connection area. This significantly improves the accuracy of the system's judgment in complex railway law enforcement scenarios, enabling the system to accurately distinguish between ordinary congestion, escalating conflict, and police safety risks, avoiding misjudgments or omissions caused by relying on a single image.

[0117] In one embodiment, a multimodal, multi-agent law enforcement assistance system for railway police chest-worn body cameras includes:

[0118] The multimodal data acquisition module is used to collect multimodal raw perception information from railway law enforcement sites and perform time alignment and scene normalization processing to generate on-site situation information;

[0119] The railway law enforcement situation awareness intelligent agent module is used to conduct a comprehensive risk assessment based on the on-site situation information and attitude feature information, and generate risk score information and risk level information.

[0120] The key evidence identification and protection intelligent agent module is used to assess the value of evidence based on the risk score information and the risk level information, trigger key evidence protection actions, and generate key evidence marking information.

[0121] The law enforcement handling assistance intelligent agent module is used to generate handling assistance prompts based on the risk level information, the key evidence marking information, and the equipment status information.

[0122] The resource scheduling intelligent agent module is used to make resource scheduling decisions based on the risk level information, the key evidence marking information, and the equipment resource status information, and to generate resource allocation instruction information.

[0123] The conflict arbitration module is used to uniformly organize and arbitrate the action request information generated by each intelligent agent module, and generate the final execution action information.

[0124] The evidence metadata management module is used to write risk assessment information, key evidence marking information, resource allocation instruction information, conflict arbitration process information, and police confirmation result information into the evidence metadata information, forming a traceable closed loop of law enforcement evidence.

[0125] A specific embodiment of the multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders is as follows:

[0126] First, the system continuously collects on-site video, audio, posture, location, equipment status, and railway business information using body cameras. This multi-source data undergoes time synchronization and scene normalization, unifying previously fragmented video footage, ambient sounds, police officer posture, location, and equipment status into a single judgment framework, generating comprehensive on-site situation information reflecting the current law enforcement situation. Next, acceleration and angular velocity features are extracted from the on-site situation information to generate posture feature information. Then, multi-dimensional information such as video anomalies, audio anomalies, location sensitivity, equipment anomalies, and posture features are weighted and fused to generate a risk score, which is compared with a preset threshold to determine the risk level. This allows the system to accurately distinguish the urgency of different events such as arguments in train carriages, pushing on platforms, police officer falls, and equipment obstruction. Then, based on risk level, event type, audio-visual integrity, and railway scene location, a comprehensive assessment of evidence value is performed. When the score reaches a trigger threshold, the system automatically locks the audio-visual cache within a preset time period before and after the event, generating key evidence markers containing the event's time, location, and type, and performing anti-overwrite protection and integrity verification. Simultaneously, based on risk level, evidence markers, police officer posture, and equipment status... The system generates situational awareness and evidence protection prompts, including different types and intensities of prompts. It also adaptively selects low-interference prompts such as vibration, indicator lights, text, or voice based on ambient noise and lighting conditions. Simultaneously, it determines task priorities based on risk levels and evidence markers, and generates resource allocation instructions such as video quality adjustment, storage allocation, upload scheduling, and prompt channel allocation, taking into account current battery power, storage, network, and computing resources. When the four intelligent agents—situational awareness, evidence protection, handling assistance, and resource scheduling—make different requests for the same resource (camera, storage, network, or prompt channel), the system queries the resource control table and makes a unified decision based on arbitration rules prioritizing police safety, key evidence protection, high-risk events, and low-interference prompts. When law enforcement judgments or event escalation are involved, it generates police confirmation requests and updates execution actions based on feedback. Finally, it summarizes risk scores, key evidence markers, resource allocation instructions, conflict arbitration processes, and police confirmation results to generate evidence metadata, which is then stored in association with key evidence files, verified for integrity, and jointly sealed. This forms a complete and traceable closed loop of law enforcement evidence, from on-site awareness, risk identification, evidence protection, resource coordination, police confirmation, to post-event review.

[0127] The present invention and its embodiments have been described above. This description is not restrictive. The accompanying drawings are only one embodiment of the present invention. The actual structure is not limited to this. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.

Claims

1. A multimodal, multi-agent law enforcement assistance method for railway police chest-worn body camera devices, characterized in that: Includes the following steps: S10: Collect multimodal raw perception information from the railway law enforcement site, perform time alignment and scene normalization processing on the multimodal raw perception information, and generate on-site situation information; S20: Extract attitude feature information based on the on-site situation information, and perform a comprehensive risk assessment by combining the on-site situation information and the attitude feature information to generate risk score information and risk level information; S30: Based on the risk score information and the risk level information, the evidence value is assessed. When the assessment result meets the preset conditions, a key evidence protection action is triggered to generate key evidence marking information and lock the corresponding audio and video cache information. S40: Generate handling assistance prompts based on the risk level information, the key evidence marking information, and the equipment status information, and select the corresponding prompt method according to the current on-site environment status; S50: Based on the risk level information, the key evidence marking information, and the equipment resource status information, make resource scheduling decisions and generate resource allocation instruction information; S60: The agent action request information generated in steps S10 to S50 is uniformly organized and conflict arbitration is performed to generate the final execution action information; S70: Write the risk scoring information, the key evidence marking information, the resource allocation instruction information, the conflict arbitration process information, and the police confirmation result information into the evidence metadata information to form a traceable law enforcement evidence closed loop.

2. The multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders according to claim 1, characterized in that: Step S10 specifically includes: S101: Collect on-site video information, on-site audio information, equipment posture information, location information, equipment operating status information, and railway business information through a chest-worn law enforcement recorder to generate the multimodal raw perception information; S102: Time synchronization and alignment of various types of information in the original multimodal sensing information, and filtering of abnormal data to generate synchronized multimodal sensing information; S103: Perform scene normalization processing on the synchronized multimodal perception information to generate the on-site situation information.

3. The multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders according to claim 1, characterized in that: Step S20 specifically includes: S201: Extract acceleration feature information and angular velocity feature information from the equipment attitude information in the on-site situation information, and generate the attitude feature information; S202: The video anomaly feature information, audio anomaly feature information, location sensitivity information, equipment anomaly status information and attitude feature information in the on-site situation information are weighted and fused to generate the risk score information; S203: Compare the risk score information with the preset risk level threshold information to generate the risk level information.

4. The multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders according to claim 1, characterized in that: Step S30 specifically includes: S301: Based on the aforementioned risk level information, event type information, audio and video integrity information, and railway scene location information, a comprehensive score of evidence value is generated to produce evidence value score information; S302: When the evidence value scoring information reaches the preset evidence protection trigger threshold, the audio and video cache information within a preset time period before and after the event is automatically locked, and key evidence marking information containing event time, location and event type information is generated; S303: Perform anti-overwriting protection and integrity verification on the audio and video cache information corresponding to the key evidence marking information, and generate evidence protection status information.

5. A multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders according to claim 1, characterized in that: Step S40 specifically includes: S401: Based on the risk level information, the key evidence marking information, the police officer's posture and status information, and the equipment status information, generate auxiliary prompt information for handling that includes the prompt type and prompt intensity; S402: Obtain current ambient noise intensity information and ambient lighting conditions information, and select a corresponding prompting method based on the ambient noise intensity information and ambient lighting conditions information. The prompting method includes vibration prompting, indicator light prompting, text prompting, or voice prompting.

6. The multimodal multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders according to claim 1, characterized in that: Step S50 specifically includes: S501: Obtain current device resource status information, which includes power information, storage space information, network status information, and computing load information; S502: Determine task priority information based on the risk level information and the key evidence marking information; S503: Generate resource allocation instruction information based on the task priority information and the device resource status information. The resource allocation instruction information includes recording quality adjustment instruction, storage space allocation instruction, network upload scheduling instruction, and prompt channel allocation instruction.

7. A multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders according to claim 1, characterized in that: Step S60 specifically includes: S601: Obtain action request information generated by the situational awareness agent, evidence protection agent, handling assistance agent, and resource scheduling agent respectively. The action request information includes action source information, target resource information, event level information, and priority information. S602: Based on the action request information, query the resource control table information to obtain the current resource occupancy status information and occupancy source information; S603: When multiple action request messages conflict for the same resource, the final execution action information is generated according to the preset arbitration rules. The preset arbitration rules include the police safety priority rule, the key evidence protection priority rule, the high-risk event priority rule, and the low-interference prompt priority rule. S604: When the final execution action information involves law enforcement judgment or event escalation, generate police confirmation request information and update the final execution action information based on the police confirmation feedback information.

8. A multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders according to claim 1, characterized in that: Step S70 specifically includes: S701: The risk scoring information, the key evidence marking information, the resource allocation instruction information, the conflict arbitration process information, and the police confirmation result information are summarized to generate evidence metadata information; S702: Associate the evidence metadata information with the corresponding key evidence files and perform integrity verification on the evidence metadata information and the key evidence files to generate evidence integrity verification information; S703: Seal the evidence metadata information together with the key evidence documents to form a traceable closed loop of law enforcement evidence.

9. A multimodal, multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders according to claim 1, characterized in that: The railway business information includes station information, platform information, train number information, carriage information, operating section information, and duty task type information.

10. A multimodal multi-agent law enforcement assistance system for railway police chest-worn law enforcement recorders, used to execute the multimodal multi-agent law enforcement assistance method for railway police chest-worn law enforcement recorders as described in any one of claims 1-9, characterized in that, include: The multimodal data acquisition module is used to collect multimodal raw perception information from railway law enforcement sites and perform time alignment and scene normalization processing to generate on-site situation information; The railway law enforcement situation awareness intelligent agent module is used to conduct a comprehensive risk assessment based on the on-site situation information and attitude feature information, and generate risk score information and risk level information. The key evidence identification and protection intelligent agent module is used to assess the value of evidence based on the risk score information and the risk level information, trigger key evidence protection actions, and generate key evidence marking information. The law enforcement handling assistance intelligent agent module is used to generate handling assistance prompts based on the risk level information, the key evidence marking information, and the equipment status information. The resource scheduling intelligent agent module is used to make resource scheduling decisions based on the risk level information, the key evidence marking information, and the equipment resource status information, and to generate resource allocation instruction information. The conflict arbitration module is used to uniformly organize and arbitrate the action request information generated by each intelligent agent module, and generate the final execution action information. The evidence metadata management module is used to write risk assessment information, key evidence marking information, resource allocation instruction information, conflict arbitration process information, and police confirmation result information into the evidence metadata information, forming a traceable closed loop of law enforcement evidence.