Cooperative management and control method of prison big data center based on multi-modal intelligent agent

By connecting video, audio, and environmental events in chronological order within the prison big data center, identifying and processing boundary-crossing actions, tracking behavioral trajectories, and rearranging discontinuous segments, the problem of coarse granularity in behavior recognition and action conflicts in the prison big data center is solved, and orderly multimodal intelligent agent collaborative management and control is realized.

CN121547607AInactive Publication Date: 2026-02-17SHANDONG UNIV OF POLITICAL SCI & LAW +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511683939.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies in the collaborative management of multi-source data in prison big data centers have coarse behavior recognition granularity, fixed permission boundaries, and difficulty in adapting to dynamic scenarios. This leads to delayed response after actions exceed boundaries, task conflicts and path blockages, disordered resource scheduling, lack of tracking mechanisms before and after behavior intervention, and simplistic path termination judgments, making it impossible to form a basis for behavior evolution.

Method used

By extracting behavioral fragments from video, audio, and environmental events and connecting them in chronological order, a multimodal behavioral content list is generated. Out-of-bounds actions are identified and removed, the continuity of action intersections is checked, behavioral trajectories are tracked and discontinuous segments are rearranged to ensure path connectivity, and a group scheduling behavior snapshot set is generated to achieve collaborative behavior convergence.

Benefits of technology

It achieves temporal correlation of boundary crossing response, continuous control of action intervention, traceability of behavior changes, stability of group collaboration, and balanced allocation of resources, ensuring the orderly and collaborative operation of the prison big data center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547607A_ABST
    Figure CN121547607A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of agent management and control, in particular to a collaborative management and control method of a prison big data center based on a multi-modal agent, which comprises the following steps: extracting video voice and environment events, connecting according to time to generate a behavior sequence, comparing authority areas according to a task range, and eliminating border crossing actions to form controlled blocks; and checking an action direction to find an abnormal extraction intervention point, tracking a track, rearranging an interrupted section, exporting a scheduling snapshot, and checking path connection to eliminate stacking to obtain a collaborative behavior convergence state. According to the method, behavior fragments of videos, voices and environment events are connected according to a time sequence, a behavior sequence with semantic and spatial attributes is constructed, cross comparison is carried out on instruction interval positions, border crossing actions are eliminated, and controlled blocks are generated. And performing spatial check on the action sections to identify intersection abnormity, screening intervention fragments according to time coherence, tracking a track, rearranging an intermittent sequence, outputting snapshots before and after intervention, and performing comparison, so as to keep collaborative stability and resource balance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent agent management technology, and in particular to a collaborative management method for a prison big data center based on multimodal intelligent agents. Background Technology

[0002] The field of intelligent agent management technology encompasses a technical system for the collaborative management, behavioral constraints, and task allocation of multi-agent intelligence in specific scenarios. Core components include multimodal information acquisition, state correlation modeling between intelligent agents, and dynamic decision-making and control mechanisms based on multi-source data. Its overall technical scope covers intelligent agent behavior monitoring, information fusion modeling, scene situation recognition, and decision response scheduling. The focus is on achieving unified management and logical constraints on the behavioral relationships of multiple agents through computer perception and semantic understanding, thereby ensuring the orderly collaboration and stable operation of various intelligent agents in complex environments.

[0003] One collaborative management and control method for a prison big data center based on multimodal intelligent agents refers to a multi-source information fusion management and control scheme established based on the prison data center. This scheme integrates multimodal data such as voice, video, images, text, and sensor signals to identify the status of various intelligent agents in the prison management scenario, assign tasks, and constrain behavior. For collaborative management of personnel, equipment, and the environment in the prison setting, data acquisition terminals are used to collect multi-source information. A data fusion model is used to establish cross-modal feature correspondences, and a rule-driven behavior constraint mechanism is used to achieve information association and dynamic control among multiple intelligent agents. The entire system uses the prison big data center as the core data hub to complete data access, feature parsing, and command issuance, realizing collaborative processing and command coordination of multimodal information within a unified management and control architecture.

[0004] Existing technologies, even with multi-source data collaboration, suffer from coarse-grained behavior recognition and fixed permission boundaries, making them ill-suited for dynamic scenarios and leading to delayed responses after actions exceed their limits. The lack of temporal comparison at action intersections easily results in task conflicts and path blockages. Breaks in the timeline require manual repair, impacting task continuity. Behavioral interventions lack prior and subsequent tracking mechanisms, failing to provide evidence of behavioral evolution. Path termination judgments are simplistic, failing to verify action continuity and avoid stacking risks, causing resource scheduling disorder and sluggish scenario operation. Summary of the Invention

[0005] To address the technical problems existing in the prior art, this invention provides a collaborative management and control method for a prison big data center based on multimodal intelligent agents. The technical solution is as follows:

[0006] A collaborative management and control method for a prison big data center based on multimodal intelligent agents includes the following steps:

[0007] S1: Extract video, voice input and event streams from camera module and environmental monitoring device, connect action changes, semantic expressions and spatial behaviors in chronological order, set behavior identity names according to region and behavior attributes, number and map content, retain behavior fragment sequences, and generate a multimodal behavior content list;

[0008] S2: Call the semantic expression and behavior path in the multimodal behavior content list, analyze it in combination with the task execution range, identify the target segment and permission area of ​​the instruction behavior, perform position cross-analysis, find the part that exceeds the permission edge, transmit it to the smart access control and voice command node, remove the action data that is not in the limited segment, and obtain the controlled behavior block information group.

[0009] S3: Read the action range in the controlled behavior block information group, check the spatial position of the movement direction of other intelligent agents, and if there are continuous breakpoints or synchronization anomalies at the intersection, process them, extract the parts that need to be intervened, and obtain a list of collaborative intervention trigger points.

[0010] S4: Based on the list of trigger points for coordinated intervention, read the starting position of the intervention, track the behavior trajectory, reorder the segments with separate durations or repeated sequences, uniformly push the position of the discontinuous time segments, export the changes in behavior before and after the intervention as reference content, and generate a group scheduling behavior snapshot set.

[0011] S5: Invoke the behavior paths in the snapshot set of the group's scheduled behaviors, check the continuity, analyze missing, redundant, or overlapping segments in the instruction chain, check the connectivity between the preceding start point and the subsequent task termination position, ensure that the path is continuous without behavior stacking, and obtain the convergence status of the cooperative behavior.

[0012] As a further aspect of the present invention, the multimodal behavior content list includes behavior category identifiers, semantic feature encodings, spatial location information, and time series indexes; the controlled behavior block information group includes regional boundary parameters, access control data, action constraint fields, and target segment mappings; the collaborative intervention trigger point list includes interaction node information, synchronization anomaly markers, time continuity parameters, and intervention priority factors; the group scheduling behavior snapshot set includes behavior path data, time series change records, action connection status, and intervention result identifiers; and the collaborative behavior convergence status includes path connectivity indicators, behavior consistency parameters, task completion evaluation, and execution chain stability.

[0013] As a further aspect of the present invention, the steps for obtaining the multimodal behavior content list are as follows:

[0014] S101: Acquire the video frame sequence of the camera module, the language input of the voice device and the event stream of the environmental monitoring device, align the three types of information according to time tags, and perform connection processing on the interrupted segments to make the picture, voice and events correspond continuously on the same timeline to obtain multi-source time sequence segments.

[0015] S102: Extract the action changes in the picture and the semantic expression in the speech content based on the multi-source time series segments, compare adjacent actions and semantic content, and associate actions, semantics and event content according to the time sequence to make the behavior process present a continuous relationship and obtain the behavior semantic correspondence sequence.

[0016] S103: Based on the semantic correspondence sequence of the behavior, identify the spatial range involved in the behavior, divide and sort the behaviors within the same range, assign a name and number to each type of behavior, and connect the numbers sequentially into a continuous sequence to obtain a multimodal behavior content list.

[0017] As a further aspect of the present invention, the step of obtaining the controlled behavior block information group is as follows:

[0018] S201: Call the semantic expression and behavior path in the multimodal behavior content list, check the semantic items and path nodes in sequence according to the task execution scope, extract the target segment information corresponding to the action, and match each segment in sequence according to time to obtain the behavior target segment set;

[0019] S202: Based on the position comparison between the set of behavioral target segments and the boundary data of the limited permission area, detect the intersection position between the action action segment and the permission range, extract the out-of-bounds part and mark the interval boundary point to obtain the out-of-bounds segment data group;

[0020] S203: Based on the cross-boundary segment data group, input action information to the intelligent access control terminal and voice command node, perform a rejection operation on action data outside the limited range, connect the remaining segments according to the path order, and obtain the controlled behavior block information group.

[0021] As a further aspect of the present invention, the step of obtaining the list of collaborative intervention trigger points is as follows:

[0022] S301: Read the action range data in the controlled behavior block information group, compare the movement direction of each action segment with that of other intelligent agents in space, detect the location of the area where the action paths intersect, and process the intersection points in time order to obtain the action intersection interval set.

[0023] S302: Detect the action connection status at each intersection point according to the action intersection interval set, screen the segments with discontinuous action or time misalignment, extract the time data of abnormal intervals and adjacent actions and compare the start and end intervals to obtain the action connection abnormal group.

[0024] S303: Based on the abnormal action connection group, the intersection point is processed in correspondence with the adjacent instruction. The part that needs intervention is extracted based on the time interval and the action continuity value. The time node and the action direction are matched and organized to obtain a list of collaborative intervention trigger points.

[0025] As a further aspect of the present invention, the step of obtaining the group scheduling behavior snapshot set is as follows:

[0026] S401: Read the starting position of the intervention action according to the list of collaborative intervention trigger points, continuously track the behavior trajectory in the corresponding task, compare the action path in chronological order, distinguish and process the overlapping and interrupted segments, and obtain the behavior trajectory segment set.

[0027] S402: Detect the connection between actions based on the set of action trajectory segments, perform sequence adjustment on segments with duration separation or sequential repetition, redetermine the connection relationship based on the start and end times of the actions, so that the trajectory can be continuously advanced and an action connection sequence can be obtained;

[0028] S403: Based on the action connection sequence, the positions before and after the discontinuous time segments are uniformly pushed, and the action changes before and after the intervention are exported in the current time sequence to form a continuous behavior mapping, thereby obtaining a group scheduling behavior snapshot set.

[0029] As a further aspect of the present invention, the step of obtaining the convergence state of the cooperative behavior is as follows:

[0030] S501: Call the behavior path in the group scheduling behavior snapshot set, check the continuity between the start and end actions, compare the interrupted part of the path with the tail behavior, check the continuity according to the time sequence of the actions, and obtain the result path connection segment group.

[0031] S502: Detect the connection status of continuous instruction chains according to the path connection segment group, screen for missing, redundant or overlapping segments, compare the time span and sequence between actions, extract abnormal segments, and obtain a set of abnormal behavior segments;

[0032] S503: Based on the set of abnormal behavior fragments, perform connectivity checks on the starting point of the preceding behavior and the ending position of the subsequent task to determine whether the path connection remains continuous, confirm the chain where no behavior stacking occurs, and obtain the convergence state of the cooperative behavior.

[0033] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0034] In this invention, behavioral segments of video, audio, and environmental events are connected chronologically to form a behavioral sequence with semantic and spatial attributes. The positions of the instruction's effective range are cross-checked to eliminate out-of-bounds actions and generate controlled behavioral blocks. Action segments are checked spatially, identifying breakpoints and anomalies at intersections. Segments requiring intervention are selected based on temporal continuity. Behavioral trajectories are tracked, and discontinuous segments are rearranged, outputting snapshots of changes before and after intervention for comparison. Path connectivity checks ensure a consistent convergence state of the behavior. Through this processing logic, out-of-bounds responses are more temporally correlated, action interventions have continuous control, behavioral changes are traceable, group collaboration remains stable, and resource allocation is more balanced. Attached Figure Description

[0035] Figure 1 This is a flowchart of the method of the present invention;

[0036] Figure 2 This is a flowchart illustrating the process of obtaining the multimodal behavior content list of the present invention.

[0037] Figure 3 This is a flowchart illustrating the process of obtaining the controlled behavior block information group according to the present invention.

[0038] Figure 4 This is a flowchart illustrating the process of obtaining the list of collaborative intervention trigger points according to the present invention.

[0039] Figure 5 This is a flowchart illustrating the process of obtaining a snapshot set of group scheduling behavior according to the present invention.

[0040] Figure 6 This is a flowchart of the process for obtaining the convergence state of the cooperative behavior in this invention. Detailed Implementation

[0041] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0042] refer to Figures 1 to 6 A collaborative management and control method for a prison big data center based on multimodal intelligent agents includes the following steps:

[0043] S1: Extract video footage from the camera module, language input from the voice device, and event streams from the environmental monitoring device. Connect the action changes, semantic expressions, and spatial behavior sequences in chronological order. Set behavior identity names according to the areas and behavior attributes involved. After numbering and mapping the content, retain the behavior fragment sequence for retrieval to obtain a multimodal behavior content list.

[0044] S2: Call the semantic expression and behavior path in the multimodal behavior content list, check them sequentially against the task execution range, perform position cross-analysis on the target segment where the instruction behavior is applied and the limited permission area, find the part that exceeds the permission edge, input action information into the smart access control terminal and voice command node, remove the action data that is not in the limited segment, and obtain the controlled behavior block information group.

[0045] S3: Read the action range of each group in the controlled behavior block information group, check the spatial position of each action segment and the movement direction of other intelligent agents, and when it is found that there are action continuity breaks or synchronization abnormal segments at the intersection of behaviors, process the intersection point with the adjacent instructions, extract the part that needs to be intervened by the start and end interval and the temporal continuity between actions, and obtain a list of collaborative intervention trigger points.

[0046] S4: Read the starting position of the intervention action based on the list of collaborative intervention trigger points, continuously track the behavior trajectory content in the corresponding task, reorder the segments with time separation or sequential repetition between actions, uniformly push the positions of the discontinuous time segments, export the changes in behavior before and after the intervention as reference content in the current time sequence, and obtain a snapshot set of group scheduling behavior.

[0047] S5: Call the content of each behavior path in the group scheduling behavior snapshot set, check the continuity between its start and end actions, check the interrupted parts of the path with the tail behavior, analyze the missing, redundant or overlapping behavior segments between continuous instruction chains, check the connectivity between the starting point of the preceding behavior and the ending position of the subsequent task, and after confirming that the path is continuous and there is no behavior stacking, obtain the convergence status of the cooperative behavior.

[0048] The multimodal behavior content list includes behavior category identifiers, semantic feature encodings, spatial location information, and time series indexes; the controlled behavior block information group includes regional boundary parameters, access control data, action constraint fields, and target segment mappings; the collaborative intervention trigger point list includes interaction node information, synchronization anomaly markers, time continuity parameters, and intervention priority factors; the group scheduling behavior snapshot set includes behavior path data, time series change records, action connection status, and intervention result identifiers; and the collaborative behavior convergence status includes path connectivity indicators, behavior consistency parameters, task completion evaluation, and execution chain stability.

[0049] Please see Figure 2 The steps to obtain the multimodal behavior content list are as follows:

[0050] S101: Acquire the video frame sequence of the camera module, the language input of the voice device and the event stream of the environmental monitoring device, align the three types of information according to time labels, and perform connection processing on the interrupted segments to make the picture, voice and events correspond continuously on the same timeline, and obtain the result of multi-source time series segments.

[0051] The process involves acquiring video frame sequences from a camera module, speech input from a voice device, and event streams from an environmental monitoring device. These three types of information are then aligned sequentially by time labels. Interrupted segments are then stitched together to ensure continuous correspondence between video, speech, and events on a unified timeline, resulting in a multi-source time series segment. To achieve unified modeling of multi-source inputs, a model can be defined as... ,in , , respectively represent the multimodal acquisition unit set, behavioral feature mapping function, node number, process path, and observation and search state variables. This model provides the mathematical basis for establishing time consistency of multi-source signals. First, the frame sequence of the camera module is acquired, with each frame accompanied by a timestamp. First, the system identifies the location of the data on the timeline. Second, it extracts the language sequence input from the voice device and records its timestamp. Third, it acquires the event stream generated by the environmental monitoring device (such as temperature, humidity, brightness, noise, etc.) and arranges them in chronological order. Frame-level synchronization and interpolation are performed by comparing the three types of timestamps. When there is a frame interruption in the video, the missing time point can be reconstructed through a neighboring frame interpolation algorithm, thus ensuring accurate alignment of visual, voice, and environmental events on the timeline. In this way, multimodal data is fused within a unified time frame. For example, when the environmental device detects a sudden temperature rise, the video displays a change in the instrument, and the voice module captures a "heat up" command, the three types of information are synchronously aggregated to form a multi-source time segment at the same moment.

[0052] S102: Extract the action changes in the video and the semantic expression in the audio content based on the multi-source time series segments, compare adjacent actions and semantic content, and associate actions, semantics and event content according to the time sequence to make the behavior process present a continuous relationship and obtain the result behavior semantic correspondence sequence.

[0053] This process extracts action changes from video footage and semantic expressions from audio content based on multi-source time series segments. Adjacent actions and semantic content are compared, and actions, semantics, and event content are linked chronologically to present a continuous relationship in the behavioral process, resulting in a sequence of corresponding behavioral semantics. The entire process can be abstracted as a workflow diagram. ,in Nodes represent the functional positions of multi-role intelligent agents. This represents the temporal dependency between them. By mapping video actions to semantic speech through node relationships, different modalities are logically linked. First, frame difference and skeleton features are extracted from the video sequence to identify actions; second, the speech input is recognized and segmented to obtain semantic instructions; environmental monitoring events are included as auxiliary background information for comparison. When timestamps are consistent or adjacent, the semantics of the behavior determine whether they belong to the same process. For example, if a person in the video moves from left to right and the speech recognition result is "walk to the right," then the two are linked through an edge set. Connections are made; if the voice command does not match the visual, the node is split into an independent subgraph to separate the behavior. In continuous analysis, the semantic consistency and physical time interval between adjacent behaviors are compared to identify the continuity of actions and semantic dependencies, forming a complete sequence of semantic correspondences between behaviors.

[0054] S103: Identify the spatial range involved in the behavior based on the semantic correspondence sequence of the behavior, divide and sort the behaviors within the same range, assign a name and number to each type of behavior, and connect the numbers into a continuous sequence to obtain the result multimodal behavior content list.

[0055] Based on the semantic correspondence sequence of behaviors, the spatial range involved in a behavior is identified. Behaviors within the same range are then partitioned and sorted. Each type of behavior is given a name and a number, and these numbers are concatenated into a continuous sequence to obtain a list of multimodal behaviors. To ensure consistency between space and policy, an agent policy function can be introduced. ,in This is a prompt template. For toolsets, This represents the language model instance selected through the assignment function. In each state... The system generates a policy response and determines the execution rules for the current behavior. The spatial location of each behavior is determined based on visual region segmentation; actions such as "entering a room," "picking up an item," and "leaving an area" are determined by spatial coordinate differences. When a behavior changes across regions, the system... The system automatically adjusts the type of model invoked, allowing the agent to choose between reasoning and vision modes based on scene complexity. After partitioning, each action is assigned a unique number, maintaining a continuous temporal mapping. All action numbers are arranged chronologically, forming an ordered chain of actions. If actions are spatially adjacent, an edge relationship is established to ensure continuity. For example, "walking to the door" and "opening the door" can be combined into the same continuous action. This process generates a list of multimodal actions.

[0056] Please see Figure 3 The steps for obtaining the controlled behavior block information group are as follows:

[0057] S201: Call the semantic expression and behavior path in the multimodal behavior content list, check the semantic items and path nodes in sequence according to the task execution scope, extract the target segment information corresponding to the action, and match each segment in sequence according to time to obtain the behavior target segment set;

[0058] The semantic expressions and behavior paths in the multimodal behavior content list are invoked. Based on the task execution scope, the semantic items and path nodes are viewed sequentially to extract the target segment information corresponding to the actions. These segments are then matched sequentially according to time sequence to obtain the behavior target segment set. This process is achieved through a model allocation function. Description, in which This is used to control the dynamic selection of model types. When parsing semantic items and behavioral paths, the physical segment where the action occurs is first identified by comparing the time index and spatial range, and the semantic item is bound to the path node. For example, if the action "entering the living room from the kitchen" occurs in the frame interval (100, 120), this spatial transition area is extracted as the target segment; if the corresponding semantic is "entering the room," then the two match successfully. These segments are connected sequentially in chronological order to form a complete behavioral path mapping. Each segment contains start and end frames, location coordinates, the executing agent, and semantic attribute labels, forming a structured set. In this way, the target segments of all behavioral actions are ordered on a unified time axis, laying the foundation for subsequent permission detection, path planning, and boundary violation analysis. This results in a set of behavioral target segments.

[0059] S202: Based on the position comparison between the behavior target segment set and the boundary data of the limited permission area, detect the intersection of the action action segment and the permission range, extract the out-of-bounds part and mark the interval boundary point to obtain the out-of-bounds segment data group;

[0060] Based on the location comparison between the behavioral target segment set and the boundary data of the defined permission area, the intersection of the action action segment and the permission range is detected. The out-of-bounds portion is extracted and the interval boundary points are marked, resulting in the out-of-bounds segment data group. A coordinator logic and planning triggering mechanism are employed. , This is used to determine whether the planning phase has begun. When a status output is detected... When an out-of-bounds behavior is detected, the behavior planner module is invoked to recalculate the execution path. This mechanism ensures that out-of-bounds behavior is identified and responded to immediately. During the detection process, the boundary data of the target segment and the boundary model of the permission area are acquired, and the overlapping part of the behavior area and the permission range is calculated using geometric intersection. When an out-of-bounds behavior is detected, the start and end frames and spatial coordinates of the segment are immediately extracted and marked as the interval boundary point. For example, if the behavior "from kitchen to living room" crosses the restricted area, the boundary point is recorded as the intersection at the kitchen doorway. All out-of-bounds information is aggregated to form an out-of-bounds segment data group.

[0061] S203: Based on the out-of-bounds segment data group, input action information to the intelligent access control terminal and voice command node, perform a rejection operation on the action data that is not within the limited range, connect the remaining segments according to the path order, and obtain the controlled behavior block information group.

[0062] Based on the out-of-bounds segment data group, action information is input to the intelligent access control terminal and voice command node. Action data outside the limited range is eliminated, and the remaining segments are connected according to the path order to obtain the controlled behavior block information group. The core of this process lies in the planner logic. To legitimize behavior, among which It is a structured parser, responsible for filtering and structure reconstruction. This is a model instance used to parse the input state. With supplementary information During execution, the time range and coordinates of actions within the out-of-bounds data group are first identified. By comparing with the boundaries of the permission area, it's determined which actions are outside the permitted range. For out-of-bounds portions, an immediate removal operation is performed, removing the action data segment from the path sequence and generating a removal log to ensure complete traceability. For retained legal action segments, they are reconstructed and connected in chronological order to maintain the continuity of the global path. Spatially, the closure of the action area is verified to avoid path breaks caused by removal operations. For example, if the action sequence includes "opening a door → entering a room → moving out of bounds," an out-of-bounds path segment will be automatically removed, retaining only the first two legal segments, and the time mapping relationship will be recalculated to ensure natural path connection. This mechanism leverages the planner's parsing and reorganization capabilities to ensure that action blocks maintain logical integrity after removing invalid segments. A controlled action block information group is then generated.

[0063] Please see Figure 4 The steps to obtain the list of collaborative intervention trigger points are as follows:

[0064] S301: Read the action range data in the controlled behavior block information group, compare the movement direction of each action segment with that of other intelligent agents in space, detect the location of the intersection of action paths, and process the intersection points in time order to obtain the action intersection interval set.

[0065] The system reads the action range data from the controlled behavior block information group, spatially compares the movement directions of each action segment with those of other agents, detects the locations where action paths intersect, and processes the intersection points in chronological order to obtain a set of action intersection intervals. In multi-agent scenarios, the supervision logic is... Define, where This is a structured model used to generate policy assignment functions under constraints. By calculating the spatial trajectory and action time interval of each agent, its intersection points with other agents are analyzed. When two or more action paths overlap in time and space, the center position and time label of their intersection region are calculated. These intersection points are recorded as key nodes. During this process, the path function is used... The system determines the priority and behavioral dependencies of each intersection point. For example, if two individuals move simultaneously in a narrow passage, their intersection risk is calculated based on timestamps and path directions, and the order of actions is adjusted using a semantic routing mechanism. All detected intersection points are sorted chronologically and marked as follows: This is used for subsequent timing optimization and path replanning. Through this supervision mechanism, intersection risks can be automatically detected and action intersection interval sets can be generated in complex multi-body behaviors.

[0066] S302: Based on the action intersection interval set, detect the action connection at each intersection point, screen the segments with discontinuous action or time misalignment, extract the time data of abnormal intervals and adjacent actions, and compare the start and end intervals to obtain the action connection abnormal group.

[0067] Based on the action intersection interval set, the action continuity at each intersection point is detected. Segments with discontinuous actions or time misalignments are screened, and the time data of abnormal intervals and adjacent actions are extracted and compared with the start and end intervals to obtain groups of abnormal action continuity. State transition equations are then used. The temporal evolution of modeling actions, where This function aggregates input states and policy outputs. It records the time dependency of action states and policy feedback, enabling analysis of continuity. The specific process includes: first, reading the time series within the intersection interval set and extracting the action segments before and after each intersection point; second, calculating the time difference and spatial distance to identify continuity interruptions or offsets. For example, if "picking up an item" is completed within T1–T5, but "placing an item" begins at T8, the time gap between T5–T8 is detected as a discontinuous segment, triggering an anomaly flag. For cases of actions being advanced or delayed, the actual and theoretical continuity times are compared to determine if they exceed the allowable threshold. All abnormal intervals are recorded by time index, forming an action continuity anomaly group.

[0068] S303: Based on the abnormal action connection group, the intersection point is matched with the adjacent instruction and the part that needs intervention is extracted based on the time interval and the action continuity value. The time node and the action direction are matched and organized to obtain the list of collaborative intervention trigger points.

[0069] Based on the action continuity anomaly group, the intersection points are mapped to adjacent instructions for processing. The parts requiring intervention are extracted based on time intervals and action continuity values. Time nodes are then mapped to action directions to obtain a list of coordinated intervention trigger points. To make the intervention calculable, it is first formalized using a planning mapping: ,in For sub-tasks, For dependence, For resource parameters; then use supervised routing to determine the intervention decision: Read the exception group and calculate the actions before and after each intersection point. The system uses continuous values ​​and direction vectors to determine if a threshold has been crossed; if a threshold has been crossed, intervention candidates are identified. Selection is based on minimizing the global cost. Under the premise of ensuring quality constraints This simultaneously suppresses frequent switching and redundant execution. The selected time period is a set of time nodes. Output the data and archive it along with the action direction, dependencies, and resource tags for later insertion, correction, or rearrangement. Obtain a list of collaborative intervention trigger points.

[0070] Please see Figure 5 The steps for obtaining the snapshot set of group scheduling behavior are as follows:

[0071] S401: Read the starting position of the intervention action based on the list of collaborative intervention trigger points, continuously track the behavior trajectory in the corresponding task, compare the action path in chronological order, distinguish and process the overlapping and interrupted segments, and obtain the behavior trajectory segment set.

[0072] Based on the list of collaborative intervention trigger points, the starting position of the intervention action is read, and the behavioral trajectory in the corresponding task is continuously tracked. The action paths are compared in chronological order, and overlapping and interrupted segments are distinguished and processed to obtain a set of behavioral trajectory segments. This trajectory tracking follows the operator... ,in This represents the trajectory update mechanism for state evolution. After reading the list of trigger points, it updates the trajectory from the start time of each intervention point. The system traces the corresponding action path, recording the position sequence and action labels. It normalizes and compares the time differences between adjacent trajectories to determine their continuity or interruption. When a path intersection is detected, a conflict zone is identified using a spatial distance threshold and direction vector. If a time interruption is found, the missing segment is re-inserted to complete the timing sequence. The integrity of each path is determined by a function. The calculation yields a set of trajectory segments. Each segment includes time, direction, speed, and semantic identifiers. For example, in the sequence of "walking towards the door—opening the door—entering the room," if the "opening the door" time is advanced, the action is adjusted to a reasonable range through trajectory correction. Through this process, all trajectory segments form a continuous mapping on the time axis, resulting in a structured set of behavioral trajectory segments.

[0073] S402: Detect the connection between actions based on the action trajectory segment set, perform sequence adjustment on segments with duration separation or sequential repetition, redetermine the connection relationship based on the start and end times of the actions, so that the trajectory can be continuously advanced and an action connection sequence can be obtained;

[0074] Based on the behavior trajectory segment set, the connection between actions is detected. For segments with duration separation or sequential repetition, the order is adjusted, and the connection relationship is re-determined according to the start and end times of the actions to ensure continuous trajectory progression, resulting in an action connection sequence. The coordination logic follows the formula... and conditional redirection mechanism This dynamically controls the execution of the plan. When the coordinator output signal is 1, the planning node is entered to recalculate the behavior sequence. First, all action segments are extracted from the behavior trajectory segment set, sorted according to timestamps, and the duration difference between adjacent segments is calculated. If excessively long intervals or overlaps are found, they are marked as "duration separation" or "sequential repetition." The strategy function is then adjusted accordingly. The execution order of these segments is reordered to ensure a seamless transition between the end time of each action and the start time of the next action. For example, if "opening the door" ends at T5 and "entering the room" begins at T10, T10 can be adjusted to T6 at the planning layer to achieve temporal continuity. Simultaneously, spatial conflicts are detected to ensure that the adjustment does not cause behavioral logic errors. All corrected action sequences are recoded in a time-progressive manner to form an optimized action sequence.

[0075] S403: Based on the action connection sequence, the positions before and after the discontinuous time segments are uniformly pushed, and the action changes before and after the intervention are exported in the current time sequence to form a continuous behavior mapping, thus obtaining a group scheduling behavior snapshot set.

[0076] Based on the action sequence, the positions before and after time-discontinuous segments are uniformly pushed, and the changes in actions before and after intervention are exported sequentially according to the current time order, forming a continuous behavior mapping and obtaining a snapshot set of group scheduling behavior. During this process, the action sequence is first scanned to identify all segments with time discontinuities and determine their positions in the overall behavior flow. A time discontinuity refers to a time interval exceeding an allowed threshold between two actions, or an execution delay caused by external intervention. To ensure smooth connection between discontinuous segments and their preceding and following actions, model instantiation rules are used.

[0077]

[0078] in, This represents the model kernel configuration function, used to configure the model based on the current action. With state The system selects the most suitable instance from the model set based on the model's attributes to push the action time. This process can be viewed as dynamically replacing or updating the model in a multi-agent action sequence to maintain a continuous mapping of action time. During the push phase, the system first compares the preceding and following boundary frames of the discontinuous segments to determine the amount of time difference that needs adjustment; then, it uses the instantiated model... The new time nodes are calculated to ensure that the sequence of actions after the push is continuous on the global timeline. If the end time of an action such as "walking to the door" is T5, and the original start time of the subsequent action "opening the door" is T10, the model pushes the action to the T6–T8 interval based on the historical behavior rate and temporal characteristics, thereby eliminating the gap. This operation ensures logical consistency between actions at the semantic level and achieves temporal smoothing at the structural level. The changes in actions before and after the intervention are exported temporally, and the adjusted segments are recorded according to the current timeline to form a continuous behavior mapping with global consistency. Through this process, all temporal discontinuities are corrected, and the behavior trajectories after the intervention are seamlessly connected, generating a snapshot set of group scheduling behaviors.

[0079] Please see Figure 6 The steps to obtain the convergence state of cooperative behavior are as follows:

[0080] S501: Call the behavior path in the group scheduling behavior snapshot set, check the continuity between the start and end actions, compare the interrupted part of the path with the tail behavior, check the continuity according to the time sequence of the actions, and obtain the path connection segment group.

[0081] The behavior paths in the group scheduling behavior snapshot set are invoked, and the continuity between the start and end actions is examined. The interrupted parts of the path are compared with the tail-end behavior, and the continuity is checked based on the temporal sequence of the actions to obtain the path connection segment group. To ensure spatiotemporal mapping consistency, a tool operator is used. Describes the state space consisting of actions This results in a transition. Using this operator, paths can be mapped segment by segment from the snapshot set, and continuity can be analyzed. Specifically, the start and end frames and action identifiers of each path are extracted, and the time difference between adjacent actions is calculated. If the difference exceeds the threshold, it is marked as an interrupted segment. Afterwards, based on... To restore the continuity of the operation state, the breakpoint state is... Mapped to new state For example, "walking to the door" ends at T5, and "opening the door" begins at T9, generating a virtual interpolation segment T6–T8 for a smooth transition. Through dual temporal and spatial checks, all paths are reordered, connected, and uniformly recorded in the path connection segment group.

[0082] S502: Detect the connection status of continuous instruction chains based on path connection segment groups, screen for missing, redundant or overlapping segments, compare the time span and sequence between actions, extract abnormal segments, and obtain a set of abnormal behavior segments.

[0083] Based on the path connection segment group, the connection status of continuous instruction chains is detected. Segments with missing, redundant, or overlapping elements are screened, and abnormal segments are extracted by comparing the time span and sequence of actions, resulting in a set of abnormal behavior segments. This is achieved through supervised routing functions. Implement dynamic verification, where Combining the current status with the planned path Automatic constraint feedback is generated. First, all connecting segments are scanned to identify unreasonable gaps or overlaps in time, and the span value of each segment is calculated. If a segment is found to lack a preceding action or have overlapping start and end times, it is determined to be missing or overlapping. For repeated execution of the same instruction, the semantic labels are checked for consistency; if they are repeated, it is classified as a redundant segment. For example, if the action "open the door" occurs twice consecutively with overlapping times, one of the segments is marked as abnormal. All abnormal segments, along with their timestamps, are output to the behavior abnormal segment set.

[0084] S503: Based on the set of abnormal behavior fragments, perform connectivity checks on the starting point of the preceding behavior and the ending position of the subsequent task to determine whether the path connection remains continuous, confirm the chain where no behavior stacking occurs, and obtain the convergence status of the cooperative behavior.

[0085] Based on a set of abnormal behavior fragments, connectivity checks are performed between the starting point of preceding behaviors and the ending position of subsequent tasks to determine whether path connections remain continuous. Chains without behavior stacking are confirmed to obtain the convergence state of collaborative behaviors. The kernel is assigned based on the model. ,in This defines the behavior dispatch mapping for a multimodal agent. First, it reads the set of anomalous fragments, locates the start and end times and spatial nodes of the preceding and following behaviors, and compares the integrity of the connections between them. If a temporal break is detected, it is resolved by re-invoking... Assign model roles to restore logical connectivity. For example, if there are no subsequent actions after "entering the room," the basic model will complete the "closing the door" behavior to ensure path closure. Check for stacking, i.e., multiple behaviors executing concurrently in the same time interval. If there is no overlap and the timeline is continuous, the chain is considered to have converged. Output the convergence status of the cooperative behavior.

[0086] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1.A method for collaborative management of a prison big data center based on a multi-modal intelligent agent, characterized in that, Comprise the following steps: S1: extract video, voice input from camera module and event stream of environmental monitoring device, connect action changes, semantic expressions and spatial behaviors in time sequence, set behavior identity name according to area and behavior attribute, number mapping is carried out to content, keep behavior fragment sequence, generate multi-modal behavior content list; S2: call semantic expression and behavior path in the multi-modal behavior content list, analyze in combination with task execution range, identify target section and authority area of instruction behavior effect, carry out position cross analysis, find the part beyond the permission edge, transmit to intelligent access control and voice instruction node, remove action data not in limited section, obtain authorized behavior block information group; S3: read action range in the authorized behavior block information group, check spatial position with other intelligent body motion direction, find continuous breakpoint or synchronous anomaly at intersection, process, extract the part needing intervention, obtain collaborative intervention trigger point list; S4: read intervention starting position according to the collaborative intervention trigger point list, track behavior trajectory, reorder length separation or sequential repetition paragraphs, uniformly push the position of time discontinuous paragraphs, export the change behavior before and after intervention as reference content, generate group scheduling behavior snapshot set. 2.The method of claim 1, wherein: The multi-modal behavior content list comprises behavior category identifier, semantic feature code, spatial position information and time sequence index; the authorized behavior block information group comprises area boundary parameter, authority control data, action constraint field and target section mapping; the collaborative intervention trigger point list comprises interaction node information, synchronous anomaly mark, time continuity parameter and intervention priority factor; the group scheduling behavior snapshot set comprises behavior path data, time sequence change record, action connection state and intervention result identifier. 3.The method of claim 1, wherein: The acquisition step of the multi-modal behavior content list is: S101: acquire video frame sequence of camera module, language input of voice device and event stream of environmental monitoring device, align three kinds of information in time label in turn, carry out connection processing to interrupted fragment, make picture, voice and event correspond continuously on unified time line, obtain multi-source time sequence fragment; S102: extract action change in picture and semantic expression in voice content according to the multi-source time sequence fragment, compare adjacent action and semantic content, associate action, semantic and event content according to time sequence, make behavior process present continuous relationship, obtain behavior semantic corresponding sequence; S103: identify spatial range involved by behavior according to the behavior semantic corresponding sequence, sort behaviors in the same range, set name and give number to each kind of behavior, connect numbers in turn into continuous sequence, obtain multi-modal behavior content list. 4.The method of claim 1, wherein: The acquisition step of the authorized behavior block information group is: S201: call semantic expression and behavior path in the multi-modal behavior content list, sequentially view semantic item and path node according to task execution range, extract target section information corresponding to action, correspond each section in time sequence, obtain behavior target section set; S202: According to the boundary data of the behavior target section set and the defined permission area, the intersection position of the action section and the permission range is detected, the out-of-bound part is extracted and marked, and the out-of-bound section data set is obtained; S203: Based on the out-of-bound section data set, the action information is input to the intelligent access terminal and the voice instruction node, the action data outside the defined range is executed, the remaining fragments are connected according to the path order, and the authorized behavior block information set is obtained. 5.The method of claim 1, wherein: The acquisition step of the collaborative intervention trigger point list is: S301: Read the action range data in the authorized behavior block information set, compare the action path intersection with the motion direction of other intelligent agents in space, detect the area position of the action path intersection, and process the intersection points in time sequence to obtain the action intersection interval set; S302: According to the action intersection interval set, the action connection condition at each intersection point is detected, the fragments with action discontinuity or time misalignment are screened, the abnormal interval and the time data of adjacent actions are extracted and compared, and the action connection abnormal group is obtained; S303: According to the action connection abnormal group, the intersection point and the adjacent instruction are processed, the part needing intervention is extracted according to the time interval and the action continuity value, the time node and the action direction are correspondingly arranged, and the collaborative intervention trigger point list is obtained. 6.The method of claim 1, wherein the method further comprises: The acquisition step of the group scheduling behavior snapshot set is: S401: According to the collaborative intervention trigger point list, the starting position of the intervention action is read, the behavior trajectory in the corresponding task is continuously tracked, the action path is compared in time sequence, and the intersection and interruption fragments are processed to obtain the behavior trajectory section set; S402: According to the behavior trajectory section set, the connection between actions is detected, the fragments with time separation or sequence repetition are sequentially adjusted, the connection relationship is re-determined according to the start and end time of the action, the trajectory is continuously pushed forward, and the action connection sequence is obtained; S403: Based on the action connection sequence, the positions before and after the time discontinuous fragment are uniformly pushed, the action change content before and after the intervention is sequentially exported according to the current time sequence, the continuous behavior mapping is formed, and the group scheduling behavior snapshot set is obtained. 7.The method of claim 1, wherein, Further comprising: S5: Call the behavior path in the group scheduling behavior snapshot set, check the connection, analyze the missing, redundant or overlapping fragments in the instruction chain, check the connectivity of the previous starting point and the subsequent task termination position, ensure the continuous path without behavior stacking, and obtain the collaborative behavior convergence state; The collaborative behavior convergence state includes path connectivity index, behavior consistency parameter, task completion degree evaluation and execution chain stability. 8.The method of claim 7, wherein the method further comprises: The acquisition step of the collaborative behavior convergence state is: S501: Call the behavior path in the group scheduling behavior snapshot set, check the connection between the start and end actions, correspondingly compare the path interruption part with the tail behavior, check the continuity according to the time sequence between actions, and obtain the result path connection section group; S502: According to the path connection section group, the connection state of the continuous instruction chain is detected, the segments with missing, redundancy or overlap are screened, the time span and sequence between actions are compared, the abnormal paragraphs are extracted, and the behavior abnormal segment set is obtained; S503: Based on the behavior abnormal segment set, the connection of the previous behavior starting point and the subsequent task termination position is verified, whether the path connection remains continuous is judged, the chain without behavior stacking is confirmed, and the convergent state of the cooperative behavior is obtained.