Method and apparatus for track intrusion awareness based on relative optimization of group strategy
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2026-06-23
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]本申请提供一种基于群体策略相对优化的轨道入侵感知方法和装置,用于解决如何在复杂的轨道环境中对入侵威胁进行精确地感知的技术问题
[0017] The rail intrusion detection method and apparatus based on swarm strategy relative optimization provided in this application generate simulated intrusion samples by overlaying intrusion object instances into rail transit scene images. Feature annotation of the intrusion objects in the simulated intrusion samples enables the low-cost, large-scale, and high-efficiency generation of massive training data, improving the model's deep situational awareness and threat assessment capabilities. Reinforcement learning training based on the swarm strategy relative optimization algorithm ensures that the final trained rail intrusion detection model not only detects intrusion behavior but also deeply understands the context of the intrusion scene, comprehensively analyzes the spatial location and motion state of the intrusion object, and outputs accurate threat assessment results. This achieves accurate perception of intrusion threats in complex rail environments, significantly improving the intelligence level, safety, and operational efficiency of rail transit systems.
Smart Images

Figure CN122528993A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of rail transit safety monitoring technology, and in particular to a rail intrusion detection method and device based on group strategy relative optimization. Background Technology
[0002] With the rapid development of urban rail transit, track intrusion incidents pose a significant challenge to traffic safety. Traditional track intrusion detection systems mostly rely on fixed rules or sensor-based detection methods, typically only providing warnings after an incident occurs. They lack real-time assessment and early warning mechanisms for potential threats and are often unable to cope with complex and dynamic environmental changes.
[0003] Therefore, how to accurately detect intrusion threats in complex orbital environments has become a technical problem that the industry urgently needs to solve. Summary of the Invention
[0004] This application provides a method and apparatus for orbital intrusion detection based on relative optimization of group strategy, which is used to solve the technical problem of how to accurately detect intrusion threats in complex orbital environments.
[0005] This application provides a method for detecting orbital intrusions based on relative optimization of a group strategy, including: Intrusion object instances are overlaid in rail transit scene images to generate simulated intrusion samples. Feature annotations are performed on the intrusion objects in the simulated intrusion samples. The features include at least one of object position, object motion state, and threat level. The labeled simulated intrusion samples are input into the multimodal large model to be trained, and reinforcement learning training is performed based on the group policy relative optimization algorithm to generate an orbital intrusion perception model. The reinforcement learning training includes: performing multiple independent samplings on the same simulated intrusion sample to generate sampling groups, scoring each response in the sampling group, ranking the relative performance within the group based on the scores and calculating the advantage value, and updating the model parameters of the multimodal large model based on the advantage value. The real-time image of the rail transit is input into the rail intrusion perception model, and at least one of the following is obtained from the rail intrusion perception model: the predicted position of the intruding object in the real-time image of the rail transit, the predicted motion state, and the predicted threat level.
[0006] In some embodiments, prior to overlaying the intrusive object instance onto the rail transit scene image, the method further includes: Obtain raw images from at least two rail transit image datasets; Images containing both rail transit equipment and intruding objects are selected from the original images to serve as the rail transit scene images.
[0007] In some embodiments, overlaying intrusion object instances onto a rail transit scene image includes: Object instances of a preset category are extracted from the instance segmentation dataset as the intrusion object instances; the preset category includes at least one of pedestrians, animals, vehicles, and non-biological obstacles; The intruding object instance is enhanced based on the depth information and semantic context of the orbital scene; the enhancement process includes at least one of scale scaling, angle rotation, and spatial pose transformation. The enhanced instances of intrusive objects are overlaid onto the rail transit scene image.
[0008] In some embodiments, before inputting the labeled simulated intrusion sample into the multimodal large model to be trained, the method further includes: Convert the labeled content of the intrusion objects in the simulated intrusion sample into initial instructions; The initial instructions are semantically rewritten based on a pre-defined large language model to generate question-and-answer pairs with different expressions; The simulated intrusion samples, the labeled content of the intrusion objects in the simulated intrusion samples, and the question-answer pairs are serialized to generate structured training and testing sets.
[0009] In some embodiments, scoring each response within the sampling group includes: Based on the comparison results between the output format of each response and the preset format, the format reward for each response is determined. The accuracy reward for each response is determined based on the degree of matching between the predicted threat level of each response and the labeled threat level. Based on the reasoning process text of each response and the preset reasoning consistency evaluation rules, the reasoning consistency reward for each response is determined. The format reward, accuracy reward, and reasoning consistency reward are weighted and calculated to determine the score for each response.
[0010] In some embodiments, determining the reasoning consistency reward for each response based on the matching degree between the reasoning process text of each response and a preset reasoning consistency evaluation rule includes: The reasoning process text of each response and the preset reasoning consistency evaluation rules are input into a preset large language model. The preset large language model judges the degree of logical consistency between the reasoning process and the final threat level conclusion, and outputs the reasoning consistency reward for each response.
[0011] In some embodiments, the calculation weight of the accuracy reward is greater than the calculation weight of the inference consistency reward; the calculation weight of the inference consistency reward is greater than the calculation weight of the format reward.
[0012] In some embodiments, the step of ranking relatively within a group based on scores and calculating an advantage value includes: The scores of each response within the sampling group are sorted in descending order to obtain the ranking index of each response; The advantage value of each response is determined based on the ranking index of each response; the advantage value is equal to the difference between the ranking index of each response and the average ranking within the group, and then divided by the ratio of the preset number of responses in the sampling group minus one; the average ranking within the group is equal to the preset number of responses in the sampling group plus one and then divided by two.
[0013] In some embodiments, updating the model parameters of the multimodal large model based on the dominance value includes: For each response within the sampling group, calculate the ratio of the output probability of the new policy model to the output probability of the old policy network; The product of the ratio and the dominance value of each response is compared with the product after processing by the pruning function to determine the smaller value; The loss function value is obtained by summing the smaller values corresponding to each response, taking the average value, and subtracting the relative entropy penalty term. The model parameters of the multimodal large model are updated based on the loss function value; The pruning function is used to limit the ratio between a preset lower limit and a preset upper limit; the relative entropy penalty term is used to constrain the degree of difference between the new strategy model and the reference strategy model; the new strategy model is the current multimodal large model to be updated; and the old strategy model is the multimodal large model before the current update started.
[0014] This application provides an orbital intrusion sensing device based on relative optimization of a group strategy, comprising: The data processing module is used to overlay instances of intruding objects in rail transit scene images to generate simulated intrusion samples, and to annotate the intruding objects in the simulated intrusion samples with features; the features include at least one of object position, object motion state, and threat level. The model training module is used to input labeled simulated intrusion samples into the multimodal large model to be trained, perform reinforcement learning training based on the group policy relative optimization algorithm, and generate an orbital intrusion perception model. The reinforcement learning training includes: performing multiple independent samplings on the same simulated intrusion sample to generate sampling groups, scoring each response within the sampling group, ranking the relative performance within the group based on the scores and calculating the advantage value, and updating the model parameters of the multimodal large model based on the advantage value. The intrusion detection module is used to input real-time images of rail transit into the rail intrusion detection model and obtain at least one of the following: position prediction result, motion state prediction result, and threat level prediction result of the intruding object in the real-time images of rail transit output by the rail intrusion detection model.
[0015] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned orbital intrusion detection method based on relative optimization of group strategy.
[0016] This application provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned orbital intrusion detection method based on relative optimization of a group strategy.
[0017] The rail intrusion detection method and apparatus based on swarm strategy relative optimization provided in this application generate simulated intrusion samples by overlaying intrusion object instances into rail transit scene images. Feature annotation of the intrusion objects in the simulated intrusion samples enables the low-cost, large-scale, and high-efficiency generation of massive training data, improving the model's deep situational awareness and threat assessment capabilities. Reinforcement learning training based on the swarm strategy relative optimization algorithm ensures that the final trained rail intrusion detection model not only detects intrusion behavior but also deeply understands the context of the intrusion scene, comprehensively analyzes the spatial location and motion state of the intrusion object, and outputs accurate threat assessment results. This achieves accurate perception of intrusion threats in complex rail environments, significantly improving the intelligence level, safety, and operational efficiency of rail transit systems. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the orbital intrusion detection method based on relative optimization of group strategy provided in this application.
[0021] Figure 2 This is a schematic diagram of the reinforcement learning training process based on the group policy relative optimization algorithm provided in this application.
[0022] Figure 3 This is a schematic diagram of the orbital intrusion sensing device based on relative optimization of group strategy provided in this application.
[0023] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps, units, or modules is not necessarily limited to those explicitly listed, but may include other steps, units, or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0026] Most related technologies focus on track intrusion detection based on image recognition or sensor data, but these methods have limitations when dealing with changing environmental factors. For example, vision- and sensor-based detection systems cannot effectively determine complex information such as the motion state and threat level of objects.
[0027] In order to address the shortcomings of related technologies, Figure 1 This is one of the flowcharts of the orbital intrusion detection method based on relative optimization of group strategy provided in this application, such as... Figure 1 As shown, the method includes steps 110, 120 and 130.
[0028] Step 110: Overlay instances of intruding objects onto the rail transit scene image to generate simulated intrusion samples, and annotate the intruding objects in the simulated intrusion samples with features; the features include at least one of object position, object motion state, and threat level.
[0029] Specifically, the execution entity of the track intrusion detection method based on relative optimization of group strategy provided in this application embodiment is a track intrusion detection device or system. This device can be implemented in software, such as a track intrusion detection program based on relative optimization of group strategy; it can also be a device executing the track intrusion detection method based on relative optimization of group strategy, such as a terminal, computer, or server.
[0030] Rail transit scene images refer to still images or video frames acquired by image acquisition devices (such as cameras) deployed on trains, along tracks, or at stations. These images typically contain key elements of the track environment, such as rails, ballast, signal lights, and platforms. These images can come from one or more publicly available datasets or from data collected during actual operation.
[0031] Intruding object instances refer to various targets that may appear in the track area and pose a potential or actual threat to train safety. The scope of these instances can be very broad, including pedestrians, various animals (such as dogs and cows), vehicles (such as stray cars or motorcycles), and non-biological obstacles (such as fallen rocks, fallen trees, and large pieces of trash). These object instances can be extracted from large-scale instance segmentation datasets or obtained by matting real-world objects.
[0032] Generating simulated intrusion samples involves overlaying instances of the aforementioned intrusion objects onto the background of a rail transit scene image using image synthesis techniques. To make the synthesized samples more realistic, a series of spatial transformations can be performed on the intrusion object instances, such as scaling them according to the perspective of the scene, rotating them according to their possible motion postures, or placing them in a reasonable spatial location based on depth information (e.g., on the rails, ballast, or in a safe area beside the track). In this way, dangerous intrusion scene samples that are difficult to collect in the real world or that occur only occasionally can be generated at low cost and on a large scale, greatly enriching the diversity and coverage of the training data.
[0033] Feature annotation is the process of adding information to generated simulated intrusion samples, providing ground truth for subsequent supervised and reinforcement learning of the model. Annotated features include object location, object motion state, and threat level, among others.
[0034] Object location describes the spatial region where an intruding object is situated within the orbital environment. It can be categorized into several types, such as "within the track," "in the ballast area," or "in the safe zone." When labeling, a corresponding location label can be added to the bounding box of each intruding object.
[0035] The motion state of an object is used to describe the dynamic information of an intruding object. Since training data is usually static images, the motion state here can be a pre-defined, simulated label used to allow the model to learn the threat differences in different motion patterns. For example, it can be defined as "stationary," "threatening motion" (such as crossing the track or moving along the track), or "non-threatening motion" (such as moving parallel to the track in a safe zone).
[0036] Threat level is a comprehensive assessment label and a core objective that the model ultimately needs to learn and predict. This level is typically determined based on an object's location, motion, and other potential factors (such as object size and type). For example, it can be categorized into three levels: "safe," "potential threat," and "serious threat." A pedestrian stationary in a safe area might be labeled "safe," while a pedestrian crossing a track would be labeled "serious threat."
[0037] It should be noted that only one of the above features can be labeled, such as only labeling "threat level"; or any two or all three can be labeled to provide richer supervisory information.
[0038] Step 120: Input the labeled simulated intrusion samples into the multimodal large model to be trained, and perform reinforcement learning training based on the group policy relative optimization algorithm to generate the track intrusion perception model; the reinforcement learning training includes: performing multiple independent samplings on the same simulated intrusion sample to generate sampling groups, scoring each response within the sampling group, ranking the relative performance within the group based on the scores and calculating the advantage value, and updating the model parameters of the multimodal large model based on the advantage value.
[0039] Specifically, a multimodal large model to be trained refers to a deep learning model that can simultaneously understand and process multiple modalities of information (including at least images and text). Structurally, it typically includes an image encoder for extracting visual features from the input image and a language decoder for understanding textual instructions and generating textual responses.
[0040] Then, reinforcement learning training is performed based on the Group Relative Policy Optimization (GRPO) algorithm to generate the track intrusion detection model. The core idea of reinforcement learning training is to allow the model to learn an optimal policy through interaction with the environment (in this case, the evaluation task). The unique feature of the GRPO algorithm is that it does not rely on an absolute value judgment, but rather optimizes the policy through relative comparisons within a group. The reinforcement learning training specifically includes: (1) Generate sampling groups by independently sampling the same simulated intrusion sample multiple times: For a given input (e.g., a simulated intrusion image and a text instruction such as "Please assess the threat level of the target in the image"), generate sampling groups in parallel and independently using the multimodal large model to be trained under the current strategy. A different output response. Each response together constitutes a sampling group. For example, It can be set to an integer of 4, 8, or larger. The responses may differ slightly in the reasoning process or the final conclusion.
[0041] (2) Scoring each response within the sampling group: Design a reward function or scoring mechanism to score each response in the sampling group. The scoring criteria may include, but are not limited to, whether the threat level output by the model is consistent with the threat level label manually annotated, and whether the output format meets the preset requirements.
[0042] (3) Rank the relative scores within the group and calculate the advantage value: Rank the scores within the sampled group. The responses are sorted from highest to lowest according to the scores obtained in the previous step, resulting in a relative ranking for each response within its group. Then, an advantage value is calculated for each response based on this relative ranking. The core idea of the advantage value is that it measures whether a response performs better or worse relative to the average performance of other responses in its group. Responses with higher rankings receive positive advantage values, while those with lower rankings receive negative advantage values. This relative comparison method is more stable and efficient than learning an absolute value function.
[0043] (4) Updating the model parameters of the multimodal large model based on the dominance value: The calculated dominance value is used as a guiding signal to update the parameters of the multimodal large model to be trained through the backpropagation algorithm. Specifically, for responses with positive dominance values, the model will adjust its parameters to increase the probability of generating similar responses in the future; for responses with negative dominance values, the probability of generating similar responses will be reduced. By continuously repeating this process on the entire training set, the model's strategy will be continuously optimized and eventually converge.
[0044] After sufficient reinforcement learning training, the resulting multimodal large model with converged parameters is the track intrusion sensing model to be obtained in this application embodiment.
[0045] Step 130: Input the real-time image of the rail transit into the rail intrusion perception model to obtain at least one of the following: the predicted position of the intruding object in the real-time image of the rail transit output by the rail intrusion perception model, the predicted motion state, and the predicted threat level.
[0046] Specifically, after the model is trained, the system can predict the threat level of real-time data.
[0047] Real-time images of rail transit are input into the rail intrusion perception model. The real-time images of rail transit can be one or more frames from a continuous video stream captured by cameras deployed in a real environment.
[0048] After receiving real-time images, the orbital intrusion perception model performs inference analysis and outputs its prediction results in structured text (such as JSON format) or other preset formats.
[0049] The specific process includes: (1) Location assessment: Based on image data, determine whether the object is located within the track, ballast or safety zone, and determine the spatial location of the object.
[0050] (2) Motion status assessment: The threatening nature of an object is assessed based on its motion status. Threatening motion increases the threat level, while objects with slow or non-threatening motion have a lower threat level.
[0051] (3) Threat level assessment: Taking into account the object's position and motion state, as well as other environmental factors, the model outputs the object's threat level.
[0052] For example, the model's output might be: "{'Target':'Pedestrian','Location':'Within Track','Movement Status':'Threatening Motion','Threat Level':'Severe Threat'}". These outputs can be directly used in downstream early warning systems to trigger corresponding audible and visual alarms or send alert information to the operations control center, thereby providing decision support for driving safety.
[0053] The track intrusion detection method based on group policy relative optimization provided in this application generates simulated intrusion samples by overlaying intrusion object instances into track traffic scene images. Feature annotation of the intrusion objects in the simulated intrusion samples enables the low-cost, large-scale, and high-efficiency generation of massive training data, improving the model's deep situational awareness and threat assessment capabilities. Reinforcement learning training based on the group policy relative optimization algorithm ensures that the final trained track intrusion detection model not only detects intrusion behavior but also deeply understands the context of the intrusion scene, comprehensively analyzes the spatial location and motion state of the intrusion object, and outputs accurate threat assessment results. This achieves accurate perception of intrusion threats in complex track environments, significantly improving the intelligence level, safety, and operational efficiency of the track traffic system.
[0054] It should be noted that each implementation method of this application can be freely combined, rearranged, or executed individually, and does not need to rely on or depend on a fixed execution order.
[0055] In some embodiments, before overlaying instances of intruding objects onto the rail transit scene image, the method further includes: Obtain raw images from at least two rail transit image datasets; Images containing both rail transit equipment and intruding objects are selected from the original images to serve as rail transit scene images.
[0056] Specifically, a rail transit image dataset refers to a collection of images collected and labeled specifically for rail transit scenarios. Original images can be integrated from multiple different datasets.
[0057] In one specific embodiment, this application can integrate the RailSem19 Semantic Understanding Dataset for Railway Scenes and the Rail Transit Multimodal Near-Range Remote Sensing Dataset (MRSI). The RailSem19 dataset contains a large number of real images of railway scenes with rich semantic segmentation annotations. The MRSI dataset provides synthetic images in diverse subway environments. By integrating these two datasets from different sources and with different scenes, original images covering urban subways and intercity railways, various geographical environments, different lighting conditions, and different weather conditions can be obtained.
[0058] After integrating a massive amount of original images from multiple datasets, these original images need to be screened in order to improve the efficiency of data processing and the quality of subsequent synthetic samples. The screening criteria is that the images must contain two types of key elements at the same time: (1) Rail transit equipment: This is the basic background element that constitutes the rail transit scene. Specifically, it may include, but is not limited to, fixed facilities commonly found in rail transit systems such as rails, ballast, signal lights, catenary pillars, platform edges, and tunnel entrances. (2) Intruding objects: This can also be understood as objects of interest (OOI). It refers to targets that already exist in the original images and are related to the intrusion perception task of this application embodiment, such as pedestrians, animals, or vehicles that have already been photographed in the images.
[0059] Filtering images containing these targets ensures that the selected background scene itself possesses a certain degree of complexity and semantic information, more closely resembling the real-world scenario where an intrusion might occur. The images obtained after the above acquisition and filtering steps serve as rail transit scene images, used for subsequent intrusion object instance overlay operations.
[0060] The track intrusion detection method based on group strategy relative optimization provided in this application greatly enriches the source of training data by integrating at least two datasets. By screening images containing rail transit equipment and intruding objects, the efficiency and quality of subsequent data augmentation and synthesis are improved, enabling the trained track intrusion detection model to have stronger environmental adaptability and generalization ability.
[0061] In some embodiments, overlaying intrusive object instances onto a rail transit scene image includes: Extract object instances of a preset category from the instance segmentation dataset as intrusion object instances; the preset category includes at least one of pedestrians, animals, vehicles, and non-biological obstacles; The intrusion object instance is enhanced based on the depth information and semantic context of the orbit scene; the enhancement process includes at least one of scale scaling, angle rotation and spatial pose transformation. The enhanced instances of intrusive objects are overlaid onto the rail transit scene image.
[0062] Specifically, instance segmentation is a computer vision task that not only identifies the categories of objects in an image but also generates a pixel-level precise mask for each individual instance of an object. Using this type of dataset, it is very convenient to extract object instances from their original background without any extra background pixels.
[0063] In one specific embodiment, a Large Vocabulary Instance Segmentation (LVIS) dataset can be used as the source of intrusion object instances. When extracting intrusion object instances, a predefined list of categories can be defined based on the actual needs of the rail intrusion scenario. These predefined categories can include pedestrians, animals, vehicles, and non-biological obstacles. Non-biological obstacles include obstacles caused by natural factors (such as rocks) or obstacles caused by human factors (such as discarded waste). By extracting intrusion object instances from the dataset, the diversity of intrusion object instances can be ensured, covering the vast majority of target types that may pose a threat to rail transit.
[0064] A series of enhancement processes can be performed on intrusion object instances to ensure that they blend naturally and reasonably with the background scene. The enhancement processes here are mainly based on two key pieces of information: (1) Depth information of the track scene: Depth information describes the distance from each point in the scene to the camera. Based on the depth information, the intrusion object instances can be scaled in accordance with the perspective rule of "nearer objects appear larger and farther objects appear smaller". (2) Semantic context of the track scene: Semantic context refers to the understanding of the scene content, such as where the railway tracks are, where the ballast is, and where the sky is. Based on the semantic context, the intrusion object instances can be reasonably rotated and transformed in spatial pose.
[0065] Enhancement processing may include at least one or any combination of scaling, angular rotation, and spatial pose transformation, with the aim of ensuring that the visual presentation of the intrusive object instance is consistent with the physical laws and logical relationships of the background scene.
[0066] After completing the above enhancement process and determining the size, pose, and precise location of the intrusion object instance in the background image, the final overlay operation is performed to generate a realistic simulated intrusion sample containing a specific intrusion scenario.
[0067] The orbital intrusion perception method based on relative optimization of group strategy provided in this application extracts intrusion object instances from a large-scale instance segmentation dataset, ensuring the diversity of intrusion object types and the quality of instances. By combining the depth information and semantic context of the usage scenario to transform the object instances, the realism and logical rationality of the simulated intrusion samples are greatly improved, making the trained orbital intrusion perception model have stronger environmental adaptability and generalization ability.
[0068] In some embodiments, before inputting the labeled simulated intrusion samples into the multimodal large model to be trained, the method further includes: Convert the annotations of the intrusion objects in the simulated intrusion sample into initial instructions; Based on a pre-defined large language model, the initial instructions are semantically rewritten to generate question-and-answer pairs with different expressions; The simulated intrusion samples, the labeled content of the intrusion objects in the simulated intrusion samples, and the question-answer pairs are serialized to generate structured training and test sets.
[0069] Specifically, after labeling the features of objects in the simulated intrusion samples (such as object location, motion state, and threat level), this structured labeled data needs to be converted into natural language format so that large language models can understand it. This process can be accomplished using a pre-defined core template. This template can string together discrete label information into a logically coherent and complete descriptive text sentence or paragraph.
[0070] For example, suppose an intruding object in a simulated intrusion sample is labeled as follows: object category is "pedestrian", bounding box coordinates are... The object's location is "within the track," its motion state is "threatening motion," and its threat level is "serious threat." A core template can be designed as follows: "Analysis is performed on the [category] object at coordinate [coordinate] in the image. Analysis result: The object is located at [position], its motion state is [motion state], therefore, its threat level is determined to be [threat level]." After filling the above annotations into this template, an initial command or initial response text will be generated: "Analysis is performed on the [category] object at coordinate [coordinate] in the image." The pedestrian object was analyzed. Analysis results: The object is located within the track, and its motion is threatening; therefore, its threat level is assessed as serious. This initial instruction forms the basis for subsequent processing.
[0071] A pre-defined large language model can be used to perform diverse semantic rewriting of initial instructions. This pre-defined large language model can be a general-purpose, pre-trained language model with powerful natural language understanding and generation capabilities. When performing the rewriting task, the initial instruction can be input into the model, along with the rewriting instructions, such as "Please restate the following information using various sentence structures and question formats, keeping the core semantics unchanged." Based on the instructions, the pre-defined large language model will generate a large number of question-answer pairs with the same semantics but different expressions.
[0072] Through diverse semantic rewriting, a single labeled sample is expanded into multiple training instances containing different question styles and answer methods, greatly enriching the diversity of training data.
[0073] To facilitate reading and processing by computer programs, all the aforementioned information needs to be integrated and encapsulated into a unified, structured file format. In a specific embodiment, JSON (JavaScript Object Notation) file format is used for encapsulation.
[0074] Each piece of structured data can contain the following information: the image path of the original simulated intrusion sample, the original annotation content of the intrusion object (such as bounding box coordinates, category, location, etc.), and one or more question-answer pairs generated by a large language model.
[0075] Finally, all generated structured data files are divided into training and test sets according to a preset ratio. The training set is used for model parameter learning and optimization, while the test set is used to evaluate the model's performance and generalization ability after training.
[0076] The orbital intrusion perception method based on relative optimization of group strategy provided in this application converts structured annotations into natural language question-answer pairs, enabling the data format to seamlessly integrate with the training paradigm of multimodal large models and improving the effectiveness of training. By using large language models for semantic rewriting, the diversity of language expression is greatly enriched, effectively avoiding overfitting of the model to specific text templates, and enabling the model to learn the true relationship between visual features and high-level semantics, thereby significantly enhancing the robustness and generalization ability of the final generated orbital intrusion perception model.
[0077] In some embodiments, scoring is performed on each response within a sampling group, including: Based on the comparison results between the output format of each response and the preset format, the format reward for each response is determined. The accuracy reward for each response is determined based on the degree of matching between the predicted threat level of each response and the labeled threat level. Based on the reasoning process text of each response and the preset reasoning consistency evaluation rules, the reasoning consistency reward for each response is determined. The format reward, accuracy reward, and reasoning consistency reward are weighted and calculated to determine the score for each response.
[0078] Specifically, format rewards are used to ensure that the large multimodal model to be trained can generate structured, normalized outputs.
[0079] A standard output format needs to be predefined, such as JSON with multiple keys. During scoring, the system checks whether each response text generated by the model can be successfully parsed into the standard output format. If a response can be fully parsed and all required keys are present, it is given a high format reward value, such as 1.0. If a response has partial formatting defects, such as missing a non-critical key, it can be given a lower partial reward value, such as 0.5. If a response is entirely free text and cannot be parsed into the predefined structured format, its format reward value is 0.
[0080] By setting format rewards, the output behavior of the model can be effectively constrained, enabling it to generate machine-friendly and directly usable results.
[0081] Accuracy rewards are used to improve the accuracy of model evaluation results. During scoring, the system extracts the predicted threat level from the model response and compares it with the ground truth labels given during the data annotation phase.
[0082] In one specific implementation, different levels of reward can be assigned based on the degree of matching: if the threat level predicted by the model (e.g., "serious threat") is exactly the same as the labeled threat level, the highest accuracy reward value is given, such as 1.0. If the threat level predicted by the model deviates somewhat from the labeled level but is not completely wrong (e.g., labeled "serious threat," the model predicts "potential threat"), a partial reward value between 0 and 1 can be given, such as 0.5. This setting encourages the model to make near-correct predictions rather than completely wrong predictions (e.g., predicting "safe"). If the threat level predicted by the model is completely different from the labeled level (e.g., labeled "serious threat," the model predicts "safe"), the lowest accuracy reward value is given, typically 0.
[0083] In this way, accuracy rewards can most directly and powerfully guide the model to learn correct threat assessment capabilities.
[0084] The reasoning consistency reward is used to ensure that the model not only provides the correct answer, but also offers a logically consistent reasoning process that supports its conclusion. This is crucial for improving the interpretability and reliability of the model.
[0085] During scoring, the system first extracts the reasoning process text from the model's response. Then, it determines whether the reasoning process logically leads to the final threat level conclusion based on preset reasoning consistency evaluation rules. Evaluation rules can be a series of logical judgment conditions, such as: "Rule 1: If the object's location is within the 'track' or 'ballast area' and its movement is 'threatening movement,' then its conclusion should be 'serious threat.'" The system checks whether the model's reasoning process and conclusion conform to these preset rules. If they do, a higher reasoning consistency reward (e.g., 1.0) is given; if they do not, a lower reward (e.g., 0) is given.
[0086] After receiving the three individual rewards, they need to be combined into a final comprehensive score, which can be achieved by weighted summation.
[0087] The orbital intrusion perception method based on relative optimization of group strategy provided in this application not only focuses on the correctness of the prediction results, but also takes into account the usability of the output format and the logic of the reasoning process; so that the final trained orbital intrusion perception model has high accuracy, good structure and strong interpretability in its output results.
[0088] In some embodiments, the reasoning consistency reward for each response is determined based on the matching degree between the reasoning process text of each response and a preset reasoning consistency evaluation rule, including: Input the reasoning process text of each response and the preset reasoning consistency evaluation rules into the preset large language model. The preset large language model judges the degree of logical consistency between the reasoning process and the final threat level conclusion, and outputs the reasoning consistency reward for each response.
[0089] Specifically, the pre-defined large language model here can be a general large language model or a model that has been fine-tuned for a specific logical reasoning task.
[0090] The inference process text is natural language text extracted from a specific response generated by the multimodal large model to be trained, used to explain the basis of its judgment. The final threat level conclusion is the final prediction result extracted from the same response. The pre-defined inference consistency evaluation rules are a set of guiding principles described in natural language as evaluation benchmarks. These rules are not strict program code, but rather logical rules for the pre-defined large language model to understand and follow.
[0091] The reasoning process text for each response, along with preset reasoning consistency evaluation rules, is input into a preset large language model. The preset large language model processes the input, determines the degree of logical consistency between the reasoning process and the final threat level conclusion, and outputs a numerical score. This score is used as the reasoning consistency reward for that response.
[0092] The orbital intrusion perception method based on relative optimization of group strategy provided in this application uses a preset large language model as an intelligent evaluator to determine the reasoning consistency reward, which can make more accurate and reasonable judgments than simple rule matching, making the final generated orbital intrusion perception model more reliable in decision-making and its behavior more interpretable.
[0093] In some embodiments, the calculation weight of the accuracy reward is greater than the calculation weight of the inference consistency reward; the calculation weight of the inference consistency reward is greater than the calculation weight of the format reward.
[0094] Specifically, the accuracy reward is weighted to the highest level (e.g., 75%) because accurately predicting threat levels is the core task and ultimate goal of this application's embodiments. In the critical field of rail transit safety, any false alarm or missed alarm can lead to serious consequences. Therefore, the reward function must apply the strongest guiding signal to ensure that the multimodal large model to be trained focuses the vast majority of its attention and optimization resources on learning how to make correct threat judgments from visual and textual information.
[0095] Setting the weight of the reasoning consistency reward to the second highest level (e.g., 20%) reflects the principle of ensuring accuracy while also highlighting the model's interpretability and reliability. By giving reasoning consistency a significant but secondary weight to accuracy, the model can be effectively guided to learn its inherent logical relationships, making its decision-making process more transparent and consistent with human cognition, thereby enhancing the system's reliability.
[0096] The calculation weight for the format reward is set to the minimum (e.g., 5%) because while the standardization of the output format is necessary for system integration, it is a relatively basic and easily learned task. Giving a small reward signal is sufficient to guide the model to follow the pre-defined output structure, while too high a weight may cause the model to focus excessively on format during optimization and neglect the more important tasks of accuracy and logical learning.
[0097] Figure 2 This is a flowchart illustrating the reinforcement learning training process based on the group policy relative optimization algorithm provided in this application, as shown below. Figure 2 As shown, the specific steps for training a multimodal large model using reinforcement learning may include: (1) Intra-group sample sampling: Simulated intrusion samples for a given orbital intrusion scenario Utilizing the current multimodal large model to be optimized Multiple independent samples are taken according to its output probability distribution, and parallel generation is performed to include... A sample group of responses .in, The preset number of responses for the sampling group, its value range can be: By obtaining multiple differentiated responses under the same input, a sample basis is provided for subsequent calculation of relative advantages within groups.
[0098] (2) Multi-dimensional reward evaluation: using a preset reward function For each response within the sampling group Perform independent scoring to obtain the raw reward value for each response. The reward function establishes a quality score for each response by comprehensively weighting and quantifying the format compliance, location prediction accuracy, and inference logic consistency of the response.
[0099] (3) Calculation of relative ranking and advantage value within the group: based on the original reward value Within the sample group The responses are sorted in descending order.
[0100] In some specific embodiments, relative ranking within a group and calculation of advantage values based on scores include: The scores of each response within the sampling group are sorted in descending order to obtain the ranking index of each response; The dominance value of each response is determined based on its ranking index. The dominance value equals the difference between the ranking index of each response and the average rank within the group, divided by the ratio of the preset number of responses in the sampling group minus one. The average rank within the group equals the preset number of responses in the sampling group plus one, divided by two. This can be expressed by the following formula: .
[0101] In the formula, For the first One response The advantage value; Indicates a response A descending sort index established based on the original reward value within the current sampling group.
[0102] (4) Policy gradient update and model optimization: based on the relative advantage value within the group A proxy loss function with relative entropy constraints is constructed, and the parameters of a multimodal large model are iteratively optimized through backpropagation algorithm to realize the evolution of the orbital intrusion threat level assessment strategy.
[0103] In some specific embodiments, updating the model parameters of a multimodal large model based on the dominance value includes: For each response within the sampling group, calculate the ratio of the output probability of the new policy model to the output probability of the old policy network; The product of the ratio and the dominance value of each response is compared with the product after processing by the clipping function to determine the smaller value; The loss function value is obtained by summing the smaller values corresponding to each response, taking the average value, and then subtracting a relative entropy penalty term. The model parameters of a large multimodal model are updated based on the loss function value; The pruning function is used to limit the ratio between a preset lower limit and a preset upper limit; the relative entropy penalty term is used to constrain the degree of difference between the new policy model and the reference policy model; the new policy model is the current multimodal large model to be updated; the old policy model is the multimodal large model before the current update started.
[0104] Relative entropy, also known as Kullback–Leibler (KL) divergence.
[0105] The loss function can be expressed by the formula: .
[0106] In the formula, The loss function is optimized based on the group relative policy. Let be the parameter vector to be optimized; For mathematical expectation; The input is a simulated intrusion sample; This represents the probability distribution output by a multimodal large model. For the new strategy model; This is the old strategy model; For reference strategy model; Here are the coefficients of the KL divergence penalty term; This is a KL divergence penalty term; This is the clipping function; This is the clipping threshold.
[0107] The apparatus provided in the embodiments of this application is described below. The apparatus described below can be referred to in correspondence with the method described above.
[0108] Figure 3 This is a schematic diagram of the orbital intrusion sensing device based on relative optimization of group strategy provided in this application, as shown below. Figure 3 As shown, the device includes: Data processing module 310 is used to overlay intrusion object instances in rail transit scene images to generate simulated intrusion samples and to annotate the intrusion objects in the simulated intrusion samples with features; the features include at least one of object position, object motion state and threat level; The model training module 320 is used to input the labeled simulated intrusion samples into the multimodal large model to be trained, and perform reinforcement learning training based on the group policy relative optimization algorithm to generate the track intrusion perception model. The reinforcement learning training includes: performing multiple independent samplings on the same simulated intrusion sample to generate sampling groups, scoring each response within the sampling group, ranking the relative performance within the group based on the scores and calculating the advantage value, and updating the model parameters of the multimodal large model based on the advantage value. The intrusion perception module 330 is used to input real-time images of rail transit into the rail intrusion perception model and obtain at least one of the following: position prediction result, motion state prediction result, and threat level prediction result of the intruding object in the real-time images of rail transit output by the rail intrusion perception model.
[0109] The track intrusion detection device based on swarm strategy relative optimization provided in this application generates simulated intrusion samples by overlaying intrusion object instances into track traffic scene images. Feature annotation of the intrusion objects in the simulated intrusion samples enables the low-cost, large-scale, and high-efficiency generation of massive training data, improving the model's deep situational awareness and threat assessment capabilities. Reinforcement learning training based on the swarm strategy relative optimization algorithm ensures that the final trained track intrusion detection model not only detects intrusion behavior but also deeply understands the context of the intrusion scene, comprehensively analyzes the spatial location and motion state of the intrusion object, and outputs accurate threat assessment results. This achieves accurate perception of intrusion threats in complex track environments, significantly improving the intelligence level, safety, and operational efficiency of the track traffic system.
[0110] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor, communications interface, and memory communicate with each other via the communications bus. The processor can invoke logical commands stored in the memory to execute the methods described in the above embodiments, for example: Intrusion object instances are overlaid on rail transit scene images to generate simulated intrusion samples. Features of the intrusion objects in the simulated intrusion samples are labeled; features include at least one of object position, object motion state, and threat level. The labeled simulated intrusion samples are input into a multimodal large-scale model to be trained, and reinforcement learning training is performed based on a group policy relative optimization algorithm to generate a rail intrusion perception model. The reinforcement learning training includes: multiple independent samplings of the same simulated intrusion sample to generate sampling groups, scoring each response within the sampling group, ranking the relative performance within the group based on the scores and calculating the dominance value, and updating the model parameters of the multimodal large-scale model based on the dominance value. Real-time rail transit images are input into the rail intrusion perception model to obtain at least one of the following prediction results: intrusion object position prediction, motion state prediction, and threat level prediction in the real-time rail transit images output by the rail intrusion perception model.
[0111] Furthermore, the logical commands in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several commands to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] The processor in the electronic device provided in this application embodiment can call logical instructions in the memory to implement the above method. Its specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effect, which will not be repeated here.
[0113] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments.
[0114] The specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effects, so it will not be repeated here.
[0115] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method described above.
[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0117] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for detecting orbital intrusion based on relative optimization of a swarm strategy, characterized in that, include: Intrusion object instances are overlaid in rail transit scene images to generate simulated intrusion samples, and feature annotations are performed on the intrusion objects in the simulated intrusion samples. The features include at least one of object position, object motion state, and threat level; The labeled simulated intrusion samples are input into the multimodal large model to be trained, and reinforcement learning training is performed based on the population policy relative optimization algorithm to generate the track intrusion perception model. The reinforcement learning training includes: generating sampling groups by independently sampling the same simulated intrusion sample multiple times, scoring each response within the sampling group, ranking the relative performance within the group based on the scores and calculating the advantage value, and updating the model parameters of the multimodal large model based on the advantage value. The real-time image of the rail transit is input into the rail intrusion perception model, and at least one of the following is obtained from the rail intrusion perception model: the predicted position of the intruding object in the real-time image of the rail transit, the predicted motion state, and the predicted threat level.
2. The orbital intrusion detection method based on relative optimization of swarm strategy according to claim 1, characterized in that, Before overlaying the intrusive object instance into the rail transit scene image, the method further includes: Obtain raw images from at least two rail transit image datasets; Images containing both rail transit equipment and intruding objects are selected from the original images to serve as the rail transit scene images.
3. The orbital intrusion detection method based on relative optimization of swarm strategy according to claim 1, characterized in that, The superimposed instances of intrusive objects in the rail transit scene image include: Object instances of a preset category are extracted from the instance segmentation dataset as the intrusion object instances; the preset category includes at least one of pedestrians, animals, vehicles, and non-biological obstacles; The intruding object instance is enhanced based on the depth information and semantic context of the orbital scene; the enhancement process includes at least one of scale scaling, angle rotation, and spatial pose transformation. The enhanced instances of intrusive objects are overlaid onto the rail transit scene image.
4. The orbital intrusion detection method based on relative optimization of swarm strategy according to claim 1, characterized in that, Before inputting the labeled simulated intrusion samples into the multimodal large model to be trained, the method further includes: Convert the labeled content of the intrusion objects in the simulated intrusion sample into initial instructions; The initial instructions are semantically rewritten based on a pre-defined large language model to generate question-and-answer pairs with different expressions; The simulated intrusion samples, the labeled content of the intrusion objects in the simulated intrusion samples, and the question-answer pairs are serialized to generate structured training and testing sets.
5. The orbital intrusion detection method based on relative optimization of swarm strategy according to claim 1, characterized in that, The scoring of each response within the sampling group includes: Based on the comparison results between the output format of each response and the preset format, the format reward for each response is determined. The accuracy reward for each response is determined based on the degree of matching between the predicted threat level of each response and the labeled threat level. Based on the reasoning process text of each response and the preset reasoning consistency evaluation rules, the reasoning consistency reward for each response is determined. The format reward, accuracy reward, and reasoning consistency reward are weighted and calculated to determine the score for each response.
6. The orbital intrusion detection method based on relative optimization of swarm strategy according to claim 5, characterized in that, The reasoning consistency reward for each response is determined based on the matching degree between the reasoning process text of each response and the preset reasoning consistency evaluation rules, including: The reasoning process text of each response and the preset reasoning consistency evaluation rules are input into a preset large language model. The preset large language model judges the degree of logical consistency between the reasoning process and the final threat level conclusion, and outputs the reasoning consistency reward for each response.
7. The orbital intrusion detection method based on relative optimization of swarm strategy according to claim 5, characterized in that, The calculation weight of the accuracy reward is greater than the calculation weight of the inference consistency reward; the calculation weight of the inference consistency reward is greater than the calculation weight of the format reward.
8. The orbital intrusion detection method based on relative optimization of swarm strategy according to claim 1, characterized in that, The process of ranking groups relative to each other based on scores and calculating advantage values includes: The scores of each response within the sampling group are sorted in descending order to obtain the ranking index of each response; The advantage value of each response is determined based on the ranking index of each response; the advantage value is equal to the difference between the ranking index of each response and the average ranking within the group, and then divided by the ratio of the preset number of responses in the sampling group minus one; the average ranking within the group is equal to the preset number of responses in the sampling group plus one and then divided by two.
9. The orbital intrusion detection method based on relative optimization of swarm strategy according to claim 1, characterized in that, The step of updating the model parameters of the multimodal large model based on the dominance value includes: For each response within the sampling group, calculate the ratio of the output probability of the new policy model to the output probability of the old policy network; The product of the ratio and the dominance value of each response is compared with the product after processing by the pruning function to determine the smaller value; The loss function value is obtained by summing the smaller values corresponding to each response, taking the average value, and subtracting the relative entropy penalty term. The model parameters of the multimodal large model are updated based on the loss function value; The pruning function is used to limit the ratio between a preset lower limit and a preset upper limit; the relative entropy penalty term is used to constrain the degree of difference between the new strategy model and the reference strategy model; the new strategy model is the current multimodal large model to be updated; and the old strategy model is the multimodal large model before the current update started.
10. A track intrusion sensing device based on relative optimization of swarm strategy, characterized in that, include: The data processing module is used to overlay instances of intruding objects in rail transit scene images to generate simulated intrusion samples, and to annotate the features of the intruding objects in the simulated intrusion samples. The features include at least one of object position, object motion state, and threat level; The model training module is used to input labeled simulated intrusion samples into the multimodal large model to be trained, and perform reinforcement learning training based on the population policy relative optimization algorithm to generate an orbital intrusion perception model. The reinforcement learning training includes: generating sampling groups by independently sampling the same simulated intrusion sample multiple times, scoring each response within the sampling group, ranking the relative performance within the group based on the scores and calculating the advantage value, and updating the model parameters of the multimodal large model based on the advantage value. The intrusion detection module is used to input real-time images of rail transit into the rail intrusion detection model and obtain at least one of the following: position prediction result, motion state prediction result, and threat level prediction result of the intruding object in the real-time images of rail transit output by the rail intrusion detection model.