A few-shot sea area illegal event feature extraction method based on simulation enhancement

By constructing a normalized feature system and using simulation enhancement methods, the robustness problem in feature extraction of illegal events in sea areas with few samples was solved, and stable and high-precision feature extraction was achieved under multiple case conditions, thus improving the robustness and reliability of feature extraction.

CN121834311BActive Publication Date: 2026-05-08WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV OF TECH
Filing Date
2026-03-11
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies lack robustness in feature extraction for illegal events in small-sample marine areas, making it difficult to effectively process key information in unstructured case texts. Furthermore, existing data augmentation methods lack the ability to objectively distinguish the true performance of different features, resulting in unstable feature extraction.

Method used

By constructing a normalized feature system, performing semantic merging and legal attribute verification, combining a resource-aware robustness assessment strategy, introducing a simulation-enhanced large model for adaptive evaluation and screening, and utilizing type divide-and-conquer and boundary arbitration mechanisms to construct a joint training sample set, the stability and reliability of feature extraction are achieved.

Benefits of technology

It improves the robustness of feature extraction for illegal events in sea areas with few samples, ensures the stability and accuracy of feature extraction under multiple case types, reduces noise risk, avoids negative transfer phenomenon, and achieves high-precision feature extraction under low resource conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834311B_ABST
    Figure CN121834311B_ABST
Patent Text Reader

Abstract

The application provides a few-sample sea area illegal event feature extraction method based on simulation enhancement, and relates to the field of data processing. Facing the application scene of scarce sea area illegal case text samples and multiple causes of heterogeneity, the normalized feature system is constructed to express business features under different causes, and a feature extraction framework based on text span is introduced on the basis of small-scale labeled samples to realize stable identification of nested and overlapping features; combined with a resource-aware robustness evaluation strategy, real performance short board features are accurately located, and simulation enhancement is performed directionally under the condition of keeping the feature system unchanged; through type division and boundary arbitration mechanism, simulation samples and artificial samples are safely integrated, model updating is completed under deterministic training constraints, thereby realizing high-precision and stable automatic extraction of sea area illegal event features under low-resource conditions. The implementation of the technical scheme facilitates to improve the robustness of few-sample sea area illegal event feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, specifically to a method for extracting features of illegal events in small-sample marine areas based on simulation enhancement. Background Technology

[0002] With the continuous improvement of comprehensive marine governance, marine illegal cases are showing significant complexity in terms of quantity, scale, types of cases, and forms of expression. Illegal sand mining, illegal fishing, illegal sewage discharge, and illegal dumping often occur intertwined in law enforcement practice. Case descriptions are mainly recorded in natural language text, characterized by arbitrary expressions, loose structure, inconsistent terminology, and a large amount of implicit business semantics. These unstructured case texts contain key information such as vessels, personnel, locations, times, illegal acts, operating methods, and substances involved, serving as an important foundation for supporting case analysis, law enforcement assessment, and intelligent supervision.

[0003] However, in practical engineering applications, due to the inherent constraints of marine data, such as high sensitivity, high annotation costs, and extremely limited publicly available samples, the number of finely labeled samples available for model training is typically very small. The distribution of feature types across different types of illegal activities exhibits significant heterogeneity and long-tail characteristics, and the expression of the same feature differs significantly across different types of cases, making it difficult to directly reuse feature extraction models for general domains. Existing text feature extraction methods based on sequence labeling or fixed classification heads generally rely on large-scale labeled data, making them highly sensitive to sample size and data partitioning. Simultaneously, existing data augmentation methods often rely on empirical rules or static indicators, lacking the objective ability to discern the shortcomings of different features' true performance, easily introducing invalid or even harmful simulation samples, causing negative transfer to existing stable features. Therefore, when the above-mentioned processing methods for feature extraction from illegal events in small-sample marine areas are applied to servers, the robustness of feature extraction is significantly reduced.

[0004] Therefore, there is an urgent need for a simulation-enhanced method for extracting features of illegal events in small-sample marine areas. Summary of the Invention

[0005] This application provides a simulation-enhanced method for feature extraction of illegal events in few-sample sea areas, which facilitates the improvement of the robustness of the server in feature extraction of illegal events in few-sample sea areas.

[0006] The first aspect of this application provides a simulation-enhanced method for feature extraction of illegal maritime incidents with few samples. The method includes: acquiring case texts of multiple types of illegal maritime incidents, and semantically merging and verifying the legal attributes of candidate features with business orientation in the case texts to construct a normalized feature system; based on the normalized feature system, selecting a set of labeled samples from each case text, and performing structural reconstruction and unified encoding processing on the labeled sample set to generate a training sample set; based on the training sample set, introducing a feature extraction framework based on text span to extract features from the case texts, obtaining feature extraction results containing multiple normalized features; and based on the feature extraction results, characterizing each normalized feature in different data partitioning conditions by introducing a resource-aware robustness evaluation strategy. The performance stability under the given conditions is assessed, and the sample distribution information of each normalized feature in the training sample set is combined to determine the enhancement priority of each normalized feature. Based on the enhancement priority, the feature to be enhanced is determined. For the feature to be enhanced, while keeping the normalized feature system unchanged, a simulation enhancement model is used to generate a simulation case text. Adaptive evaluation and filtering are performed on the simulation case text based on the labeled sample set to obtain a simulation sample set. A type divide-and-conquer and boundary arbitration mechanism is introduced between the simulation sample set and the labeled sample set to construct a joint training sample set. Under the constraints of a fixed model structure, fixed hyperparameter configuration, and deterministic training strategy, the simulation enhancement model is updated through the joint training sample set to output structured maritime illegal event feature results.

[0007] A second aspect of this application provides a simulation-enhanced feature extraction device for small-sample maritime illegal incidents. The device includes an acquisition module and a processing module. The acquisition module acquires case texts of various types of maritime illegal incidents and performs semantic merging and legal attribute verification on business-oriented candidate features in the case texts to construct a normalized feature system. The processing module selects a set of labeled samples from each case text based on the normalized feature system and performs structural reconstruction and unified encoding processing on the labeled sample set to generate a training sample set. The processing module further introduces a text-span-based feature extraction framework to extract features from the case texts based on the training sample set, obtaining feature extraction results containing multiple normalized features. The processing module also characterizes the feature extraction results by introducing a resource-aware robustness evaluation strategy. The processing module assesses the performance stability of each normalized feature under different data partitioning conditions and, combined with the sample distribution information of each normalized feature in the training sample set, determines the enhancement priority of each normalized feature. Based on this priority, it identifies the features to be enhanced. The processing module further utilizes a large-scale simulation enhancement model to generate simulated case text for the features to be enhanced, while maintaining the normalized feature system unchanged. It then performs adaptive evaluation and filtering on the simulated case text based on the labeled sample set to obtain a simulated sample set. The processing module also introduces a type-divide-and-conquer and boundary arbitration mechanism between the simulated sample set and the labeled sample set to construct a joint training sample set. Under constraints of a fixed model structure, fixed hyperparameter configuration, and a deterministic training strategy, it updates the large-scale simulation enhancement model using the joint training sample set to output structured maritime illegal event feature results.

[0008] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, and both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method described above.

[0009] A fourth aspect of this application provides a non-transitory computer-readable storage medium storing instructions that, when executed, perform the method described above.

[0010] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages:

[0011] By establishing a normalized feature system, business features with similar semantics but significant differences in expression across different case types are unified under the same semantic and legal framework, resolving the issues of inconsistent feature definitions and difficulty in alignment in multi-case scenarios. Based on this, the limited labeled samples undergo structural reconstruction and unified encoding, making small-scale data reusable and comparable in terms of structure and distribution. A text-span-based feature extraction framework is introduced, enabling the model to simultaneously identify nested and overlapping features, significantly improving its coverage of complex expressions in real cases. Through a resource-aware robustness evaluation strategy, feature performance stability is characterized under multi-data partitioning conditions. Combined with sample distribution information, real performance shortcomings can be accurately identified, avoiding blind data augmentation. The simulation augmentation process revolves around the features to be augmented, and adaptive evaluation and screening ensure that the simulation samples are consistent with real data in semantic distribution and legal context, reducing the risk of generated noise. Combined with type-based governance and boundary arbitration mechanisms, long-tail features are augmented while avoiding negative transfer to stable features. Ultimately, model updates are completed under a fixed model structure and deterministic training strategy, achieving stable, high-precision, and sustainable automatic extraction of features of maritime illegal events under low-resource, multi-case conditions. Therefore, this facilitates improved robustness in feature extraction for maritime illegal events with few samples. Attached Figure Description

[0012] Figure 1 A flowchart illustrating a simulation-enhanced method for extracting features of illegal events in small-sample maritime areas, provided in an embodiment of this application;

[0013] Figure 2 A schematic diagram of a module for a simulation-enhanced feature extraction device for illegal events in small sample areas provided in an embodiment of this application;

[0014] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0015] Explanation of reference numerals in the attached figures: 21. Acquisition module; 22. Processing module; 31. Processor; 32. Communication bus; 33. User interface; 34. Network interface; 35. Memory. Detailed Implementation

[0016] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0017] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0018] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0019] To address the aforementioned technical problems, this application provides a simulation-enhanced method for extracting features of illegal events in small-sample maritime areas, referring to... Figure 1 , Figure 1 The flowchart illustrates a simulation-enhanced method for extracting features of illegal events in small-sample maritime areas, as provided in this application embodiment. The method is applied to a server and includes steps S110 to S160, as follows:

[0020] S110. Obtain case texts of various types of maritime illegal cases, and perform semantic merging and legal attribute verification on candidate features with business orientation in the case texts in order to construct a normalized feature system.

[0021] Specifically, the server is a computing and storage entity that carries the operation of the methods in the embodiments of this application. It is used to centrally complete the processing of case text acquisition, preprocessing, feature system management, model inference and training updates in a network environment, and provide callable feature extraction services to upper-layer business systems. The server can be a single physical server, a virtualized server instance, a containerized computing node, or a cluster composed of multiple computing nodes. Its core feature is that it has a programmable processor, a read and write storage medium, and a network interface, so as to establish a data interaction channel with external systems such as the case management system, the law enforcement document system, and the case entry system, and realize the unified aggregation and back transmission of data.

[0022] When determining the case set that constitutes illegal case types in multiple types of sea areas, first establish a case set on the server side and generate a unique case identifier for each case in the case set. The case identifier is used to uniquely refer to the case throughout the entire process and avoid confusion caused by the same name. Subsequently, based on the case identifier, retrieve the case texts associated with the same identifier in the case management system, law enforcement document system, and case information entry system, and write each retrieved case text together with its text identifier, case identifier, generation time identifier, and source identifier into the case text library. Here, the text identifier is used to uniquely identify a single case text, the generation time identifier is used to mark the formation time of the document or record, and the source identifier is used to mark whether the data comes from the case management system, law enforcement document system, or case information entry system, so as to be able to locate the data source and maintain traceability during subsequent conflict handling and retrospective auditing. When performing character normalization processing, noise segment removal processing, and sentence segment splitting processing on the case texts in the case text library, character normalization processing is used to unify full-width and half-width characters, digital representation methods, common unit writing methods, and punctuation forms, thereby reducing duplicate features caused by character differences for the same semantics; noise segment removal processing is used to delete template headers, fixed format segments, page headers and footers, blank blocks, and repeated punctuation strings that are irrelevant to the case semantics but appear frequently, thereby avoiding noise dominating statistical significance; sentence segment splitting processing is used to divide the normalized case text into a sentence segment sequence according to Chinese punctuation and format boundaries, and bind a sentence segment identifier and a position identifier to each sentence segment in the sentence segment sequence. The sentence segment identifier is used to uniquely identify the sentence segment, and the position identifier is used to record the starting and ending character positions of the sentence segment in the case text, so that any subsequent candidate feature can be located in the original text to a specific text span and maintain a consistent annotation boundary. When performing word segmentation processing,词性标注处理 (should be "part-of-speech tagging processing" in English), and stop word filtering processing on the sentence segment sequence based on the normalized case text, word segmentation processing is used to split the sentence segment into a sequence of word elements. A word element is the smallest semantic unit for subsequent statistical significance and co-occurrence aggregation; part-of-speech tagging processing is used to attach a part-of-speech label to each word element, and the part-of-speech label is used to constrain the morphological legality of candidate features and reduce meaningless combinations; stop word filtering processing is used to剔除 (should be "remove" in English) high-frequency function words such as "the", "of", "in", and "and", to prevent them from entering the candidate feature set; subsequently, perform statistical significance analysis on the processed word elements for different case types respectively to form a preliminary screening candidate feature set. Statistical significance analysis is used to characterize the discriminability of a word element in a specific case type relative to other case types, so as to preferentially retain word elements or phrases that can better represent the business direction. The expression for the statistical significance score is:

[0023]

[0024] Where, represents the word element or candidate phrase, represents the case identifier, represents the word element in the case It should be noted that there seems to be an incomplete or incorrect expression "词性标注处理" in the original Chinese text. I translated it as "part-of-speech tagging processing" as best as possible, and also corrected "剔除" to "remove" for a more accurate translation. If there are any specific requirements or corrections for these parts, please let me know.The percentage or normalized frequency of occurrence of the corresponding information in all case texts. Indicates a set of causes of action. Indicates the size of the set of cause-of-cases. Indicate the cause of action The corresponding collection of case texts, This indicates that a term has appeared in at least one set of cause-of-fact texts. The score is determined by the number of causes of action. This score increases the weight of terms that are common in this cause of action but relatively rare in other causes of action, making the initial candidate feature set more focused on expressions with business orientation, and reducing the interference of irrelevant terms in subsequent synonym merging and semantic discrimination. When performing synonym merging, entity morphology merging, and cross-case co-occurrence aggregation on the initial screening candidate feature set, synonym merging is used to merge semantically equivalent or highly similar expressions such as "hidden arrangement," "hidden arrangement," and "straight arrangement" into the same candidate feature name, thereby avoiding the fragmentation of the same semantic meaning and resulting in sparse samples. Entity morphology merging is used to merge different spellings of the same entity, such as abbreviations, aliases, capitalization differences, and differences between numbers and Chinese characters in ship names, organization names, location names, and personnel names, thus merging them into a unified entity form, thereby stabilizing subsequent mapping relationships. Cross-case co-occurrence aggregation is used to statistically analyze the co-occurrence relationship between candidate features and their context words in all standardized case texts, forming a candidate feature context portrait. The candidate feature context portrait is used to save the common co-occurring word set, dependency relationship fragment set, and trigger phrase set for the candidate feature. The co-occurring word set is used to characterize the semantic neighborhood, the dependency relationship fragment set is used to characterize the syntactic relationship, and the trigger phrase set is used to characterize the business context triggering pattern. The co-occurrence strength can be characterized by point mutual information, and the expression for point mutual information is:

[0025]

[0026] in, Representing candidate features, Indicates contextual words, Indicates simultaneous occurrence within the same co-occurrence window and The probability, express The probability of occurrence in the entire text. express The co-occurrence window represents the probability of occurrence in the entire text, which is either within a sentence or a neighborhood defined by word distance. This metric highlights contextual words that are more semantically related to candidate features by comparing the actual probability of co-occurrence with the probability of co-occurrence under the independence assumption, thus making the contextual profile of candidate features more reflective of the real law enforcement context. When performing semantic discrimination processing on each candidate feature under the constraints of a pre-established top-level feature set and in conjunction with the candidate feature context profile, the top-level feature set provides the upper-level semantic skeleton of the normalized feature system. Top-level features are a set of categories abstracted from law enforcement elements, used to limit the normalized features to be attributable and interpretable, such as location, time, subject, behavior, quantity, involved materials, operating methods, and penalties. Semantic discrimination processing determines the semantic and syntactic roles of candidate features in sentences by using the set of trigger phrases and dependency relationship fragments in the candidate feature context profile, and accordingly determines the mapping relationship from candidate features to top-level features. This mapping relationship solidifies the constraint of which top-level feature category a candidate feature belongs to. When semantic space alignment is needed to improve cross-representation consistency, semantic vectors can be constructed for both candidate features and top-level features, and similarity can be calculated. The expression for similarity is:

[0027]

[0028] in, A semantic vector representing candidate features. The semantic vector representing the top-level features. Represents the vector dot product. and These represent vector norms, respectively. This similarity, by comparing the directional consistency of two semantic vectors, helps determine which top-level feature category a candidate feature is closer to. It is used in conjunction with rule evidence from trigger phrase sets and dependency relationship fragment sets to form a draft of a normalized feature system that covers multiple causes of action and is semantically consistent. When performing legal attribute verification on the normalized features in the draft of the normalized feature system based on the legal element set, the legal element set is used to depict the necessary element framework in law enforcement and legal contexts. Legal elements are abstract sets of subject elements, behavioral elements, time elements, place elements, means elements, quantity elements, permission elements, and consequence elements, used to ensure that the normalized features are not only semantically reasonable but also have legally applicable attributes. The legal attribute verification process establishes an element attribution relationship between each normalized feature and the legal element set, and verifies whether the applicable cause of action set of the normalized feature is consistent with its legal element constraints. When it is found that a normalized feature appears in a cause of action but the legal element attribution is not valid, or the same normalized feature name corresponds to different legal elements in different causes of action, the mapping relationship is corrected. The system either categorizes or splits normalized feature names to eliminate homonymous conflicts, while merging normalized features with interchangeable expressions but different names to eliminate redundancy. Finally, the normalized feature system is solidified into a feature dictionary containing feature identifiers, feature definitions, a set of applicable causes of action, a set of trigger phrases, and a set of negative example constraints. The feature identifiers uniquely identify normalized features, the feature definitions describe the semantic boundaries of normalized features, the set of applicable causes of action limits the effective causes of action for normalized features, the set of trigger phrases constrains typical triggering patterns of normalized features in text, and the set of negative example constraints lists expressions that are easily confused with normalized features but should be excluded. This ensures that the normalized feature system can be used for consistent annotation as well as for consistent interpretation of subsequent model training and inference outputs.

[0029] S120. Based on the normalized feature system, select a set of labeled samples from each case text, and perform structural reconstruction and unified encoding on the set of labeled samples to generate a set of training samples.

[0030] Specifically, when grouping case texts by cause of action and forming a feature distribution profile for the cause of action dimension, the case texts in the case text library are first categorized and stored according to the cause of action identifier, so that each case text belongs to only one cause of action group. The cause of action grouping is used to isolate the differences in language style and business elements of different causes of action during the statistical stage. Then, within each cause of action group, the occurrence of each normalized feature is retrieved one by one based on the normalized feature system. The occurrence frequency is used to represent the total number of times the normalized feature is triggered within the cause of action group, and the number of covered sentences and segments is used to represent the breadth of the range of sentences and segments covered by the normalized feature within the cause of action group. A sentence or segment is the smallest semantic paragraph unit obtained from the previous sentence and segment segmentation. The occurrence frequency and the number of covered sentences and segments corresponding to each normalized feature are bound to the cause of action identifier to form a feature distribution profile for the cause of action dimension. The feature distribution profile is used to characterize which normalized features are more common, have wider coverage, and are more representative in a certain cause of action, and to provide quantitative constraints for the selection of subsequent candidate labeled texts.

[0031] When selecting candidate labeled texts and constructing a balanced labeled sample set across case types under the constraint of feature distribution profile, text filtering is first performed within each case type group based on feature distribution profile. This ensures that the selected case texts simultaneously meet the constraints of containing multiple normalized features, semantic integrity, and locatable text structure. The inclusion of multiple normalized features ensures that a single case text can provide richer labeled information and alleviate the sparsity problem of small samples. Semantic integrity avoids missing feature boundaries or insufficient context due to only extracting fragments. Locable text structure ensures that subsequent span-level labels can stably fall on the determined start and end character positions. To avoid data distribution bias, the number of candidate labeled texts in each case type group is kept consistent, so that the labeled sample set presents a balanced structure in the case type dimension. This reduces the risk of overfitting the model training to high-resource case types and improves the consistency of cross-case type generalization.

[0032] When performing refined annotation processing on the labeled sample set based on the normalized feature system, the normalized feature name is used as the unique annotation label. The unique annotation label is used to constrain the same semantic concept to always use the same name in all samples and avoid label drift caused by synonym substitution. During annotation, span-level annotation is performed on text segments in the case text that are semantically consistent with the normalized feature name. Span-level annotation refers to annotating continuous character intervals rather than annotating a single character or a single word, thus naturally supporting the coexistence of nested and overlapping expressions in the same text. Each annotation result is bound to a text identifier, a normalized feature name, a start character position, and an end character position. The text identifier is used to trace back to the original case text, and the start character position and end character position are used to define the text span boundary and ensure that subsequent training and evaluation can be aligned on the same boundary, forming a traceable and reproducible annotation record.

[0033] When performing structural reconstruction on labeled samples to generate a flattened sample structure, the original labeled structure exported by the labeling tool or labeling storage system is first parsed. Fields irrelevant to interface display, hierarchical collapse, annotations, etc., are stripped away, leaving only the core fields directly related to training. Subsequently, each labeled sample is uniformly reconstructed into a flattened sample structure. A flattened sample structure refers to expressing sample content using a fixed set of fields and not relying on nested levels to interpret the meaning of the labels. A flattened sample structure includes at least a text identifier, a case identifier, a text content field, and a feature list field. The feature list field consists of multiple feature objects, each containing a normalized feature name, a start character position, and an end character position. This solidifies multiple normalized features in the same case text in an enumerable and indexable manner, ensuring that the subsequent training process can directly traverse the feature objects and recover each text span and its normalized feature name.

[0034] When performing unified encoding on the labeled sample set based on the flattened sample structure, a globally unique sample identifier is first generated for each labeled sample. The globally unique sample identifier embeds the case identifier and the original text identifier to ensure no conflict occurs when merging across systems and to support reverse tracing. Subsequently, a dictionary mapping is performed on the normalized feature names. Dictionary mapping refers to mapping the normalized feature names to stable internal codes, which are centrally maintained by the feature dictionary. This ensures that the same normalized feature always corresponds to a consistent internal code across different samples, different batches, and different annotators. Consistency verification is performed synchronously during the unified encoding process. Consistency verification is used to check the consistency relationship between the feature span position and the text content. At a minimum, it includes the following: the starting character position is less than the ending character position, the ending character position does not cross the boundary, the text span can be stably truncated in the text content, and the text span semantically matches the normalized feature name. If the consistency verification fails, the corresponding annotation results are backtracked for correction or removal, thereby reducing the impact of dirty annotations on the quality of the training sample set.

[0035] After completing unified encoding and consistency verification, when generating training, validation, and test sample sets, the fully labeled sample set is first evenly partitioned according to the case identification, ensuring that the proportion of samples for each case in the training, validation, and test sets remains consistent, preventing model evaluation from being dominated by a single case. Subsequently, within each case, the coverage balance of the normalized feature distribution is further constrained, ensuring that both high-frequency and long-tail normalized features appear in the training sample set and necessary test samples are retained in the validation and test sample sets. This guarantees that sufficient normalized feature patterns can be learned during training and that the true performance of each normalized feature can be stably characterized during evaluation. Finally, the training sample set is solidified into a standard training data format that can be directly read by the feature extraction framework. The standard training data format includes at least a globally unique sample identifier, a text content field, the internal encoding of the normalized features, and the corresponding start and end character positions, while maintaining naming and encoding rules consistent with the normalized feature system. This ensures that subsequent training, evaluation, enhancement, and backtracking stages operate in a closed loop under the same data semantic constraints.

[0036] S130. Based on the training sample set, a feature extraction framework based on text span is introduced to extract features from the case text, resulting in feature extraction results containing multiple normalized features.

[0037] Specifically, when uniformly encoding the case text and normalized feature names using a shared semantic encoder, the server first determines that the shared semantic encoder is a contextualized semantic representation model with fixed parameters, making the case text and normalized feature names comparable within the same semantic space. The case text undergoes character normalization and sub-word segmentation to obtain a word sequence, which is then fed into the shared semantic encoder after incorporating positional and segmentation encodings. This outputs a contextual representation sequence of the case text. Simultaneously, a feature name sequence is constructed for the normalized feature names and fed into the same shared semantic encoder to output normalized feature name representations. These normalized feature name representations are then aggregated from multiple perspectives to form a set of normalized feature semantic prototypes. The expression for the shared semantic encoder is:

[0038]

[0039]

[0040] in, This represents the initial representation matrix obtained by adding lexical embeddings, positional embeddings, and segmentation embeddings. Indicates the first The context representation matrix output by the layer. Indicates the number of network layers. , , They represent the first The query, key, and value linear transformation parameter matrix for multi-head attention layer. This represents a multi-head attention operator. Represents the feedforward network operator. The representation layer normalization operator; this structure, through interactive modeling of the full sequence context information at each layer, enables the same word to obtain different contextual representations in different contexts, thereby supporting the boundary identification and internal semantic aggregation of subsequent span semantic representations; the normalized feature semantic prototype set is used to transform each normalized feature name into a stable semantic template, which is obtained through the same shared semantic encoder to ensure alignment consistency with the semantic representation of the case text. When enumerating candidate text spans based on a preset span length in the context representation sequence and forming a candidate text span set, the server first enumerates candidate text spans in the word position space with minimum and maximum span length constraints, and binds the starting word position, ending word position, and span length identifier to each candidate text span; then, legality constraint screening is performed on the candidate text spans. The legality constraint screening does not only rely on simple rules, but also jointly quantifies punctuation boundary constraints, part-of-speech combination constraints, dependency local consistency constraints, and sentence / segment boundary constraints into a legality score, and uses the legality score and dynamic threshold to jointly determine the retention set, where the expression for the legality score is:

[0041]

[0042] in, Indicates the starting lexical position of the candidate text span. Indicates the position of the last term in the candidate text span. The sigmoid compression function is used to map scores to intervals. This represents a part-of-speech feature vector composed of the part-of-speech sequence within the span, the start and end part-of-speech pairs, and the keyword triggering pattern. This represents a punctuation feature vector composed of whether the span crosses a punctuation mark, the punctuation mark type, and the relative position of the punctuation mark to the boundary. This represents a dependency feature vector composed of the span-internal dependency edge coverage, internal dependency connectivity, and cross-boundary dependency edge ratio. This represents a segment feature vector composed of whether the span crosses a segment boundary, the number of segments covered by the span, and the position within the segment. , , , This represents the weight parameters of the corresponding feature vector. This represents the bias parameter; the legality score maps various structural constraints to a unified score, significantly reducing the proportion of meaningless spans while ensuring coverage in the candidate text span set. Furthermore, it achieves adaptive filtering intensity for different case text styles by introducing weight parameters for different constraint feature vectors. When constructing a text span semantic representation for each candidate text span in the candidate text span set, the server simultaneously extracts the span's internal semantic information, span start boundary information, span end boundary information, and span length information from the context representation sequence. It then employs multi-head span internal attention and dual affine boundary interaction to jointly construct a more expressive text span semantic representation. The expression for the span's internal semantic information is:

[0043]

[0044] in, The context indicates the first element in the sequence. Vector representation of each word element This indicates the number of heads of attention within the span. and Indicates the first The query and key transformation parameter matrix of each attention head. Indicates the first The projection dimension of each attention head. A coarse-grained conditional vector representing the span of candidate texts. It can be obtained by concatenating the initial boundary vector, the ending boundary vector, and the length embedding, followed by a linear transformation. Indicates the span of candidate texts Under the conditions The weight of each lexical unit's contribution to the semantics within the span. This represents the semantic vector within the span obtained from aggregation; boundary interactions use a double affine to characterize the pairwise relationship between the starting and ending boundaries, and its expression is:

[0045]

[0046] in, The boundary vector representing the position of the starting lexical unit. The boundary vector representing the position of the end word. This indicates that the biaffine tensor parameters are used to model second-order interactions. Represents the linear term parameter matrix, This represents the bias vector. This represents expanding a matrix or higher-order tensor into a vector. The final text span semantic representation gates and fuses the internal semantic vectors, boundary interaction vectors, and length embedding vectors to simultaneously preserve internal semantics, boundary localization, and length priors, thereby improving the separability and coexistence of nested and overlapping features. When performing semantic matching between the text span semantic representation and the normalized feature semantic prototype set to determine the normalized feature name or non-feature category, the server generates a query vector for each candidate text span and a prototype vector for each normalized feature semantic prototype. An angularly spaced metric-learning matching method is used to improve inter-class separability under small sample conditions. Simultaneously, non-feature category prototypes are introduced to accommodate spans that do not correspond to any normalized features. The matching score is constructed using cosine similarity, learnable scale, learnable bias, and class center regularization. The prediction distribution uses a spaced Softmax, expressed as:

[0047]

[0048] in, Indicates the span of candidate texts Normalized query vector, Indicates the first The normalized prototype vector of a normalized feature semantic prototype. This indicates the true normalized feature category index or non-feature category index corresponding to the span of the candidate text. Represents cosine similarity. This indicates a learnable or preset scaling parameter used to control the sharpness of the logarithmic distribution. The angular margin parameter is used to force the similarity gap between the true and false categories. This represents the matching loss for the candidate text span. This matching method applies a margin to the true category item, making the model more inclined to learn tighter intra-class aggregation and more significant inter-class separation in small samples, thereby reducing the risk of similar expressions being misclassified to neighboring features. When the candidate text span has a low matching score for all normalized feature semantic prototypes and is closer to the non-feature category prototype, the candidate text span is determined to be a non-feature category to avoid noisy spans being forcibly classified. When performing aggregation and conflict resolution on candidate text spans identified as normalized features and generating feature extraction results, the server aggregates all candidate text spans predicted as normalized features that do not belong to non-feature categories into a predicted span set, and writes it back with text identifiers, normalized feature names, start character positions, and end character positions consistent with the training sample set. Subsequently, conflict resolution is performed, which removes duplicates of the same type and corrects boundary jitter while retaining overlapping and nested text spans. An optimal retention strategy based on weighted interval graphs is adopted to achieve a balance between maximizing confidence and minimizing redundancy in the span set under the same normalized feature name, while allowing spans with different normalized feature names to overlap or nest in intervals. The selection objective of the weighted interval graph can be expressed as:

[0049]

[0050] in, This represents the set of all candidate spans in the predicted span set. This represents the final retained span subset. Indicates span The weighted score is obtained by fusing the matching probability, the legality score, and the boundary consistency score. The redundancy penalty coefficient is used to suppress high overlap redundancy between spans of the same type. This represents the intersection-union ratio of two spans over a character range. Indicates an indicator function, This represents the normalized feature name corresponding to the span. The objective is to reduce duplicate output and boundary drift by penalizing highly overlapping spans within the same normalized feature name, while not imposing the same redundant penalty on overlaps or nesting between different normalized feature names, so that overlapping text spans and nested text spans can be preserved and output simultaneously. After conflict resolution, the feature extraction result containing multiple normalized features is obtained, and each normalized feature is bound to the consistent text identifier and text span position information of the training sample set, so that the output result can directly enter the subsequent resource-aware robustness evaluation, simulation enhancement screening and joint training sample set construction process.

[0051] S140. Based on the feature extraction results, the performance stability of each normalized feature under different data partitioning conditions is characterized by introducing a resource-aware robustness evaluation strategy. Combined with the sample distribution information of each normalized feature in the training sample set, the enhancement priority of each normalized feature is determined, so as to determine the feature to be enhanced according to the enhancement priority.

[0052] Specifically, when constructing the evaluation data pool while keeping the normalized feature system and training sample set unchanged, the server first merges the training sample set and the validation sample set according to the sample identifier to obtain the evaluation data pool. The evaluation data pool retains the sample identifier, case identifier, and normalized feature annotation result of each sample. The evaluation data pool serves as a unified data foundation for subsequent multiple rounds of data partitioning and repeated evaluation. The sample identifier is used to uniquely locate a single sample to support cross-round backtracking. The case identifier is used to maintain the balance of the case dimension during the data partitioning stage. The normalized feature annotation result is used to provide a true annotation reference for each normalized feature. The normalized feature annotation result is stored in the form of a feature object set. Each feature object contains at least the normalized feature name, the start character position, and the end character position, so that subsequent feature alignment can maintain a consistent judgment boundary at the text span level. When performing multiple rounds of data partitioning on the evaluation data pool and generating multiple sets of data partitioning conditions, the server first sets a balance constraint between the number of partitioning rounds and the case dimension, ensuring that the data partitioning conditions in each round remain comparable at the case identifier level. The evaluation data pool is then stratified and sampled according to case identifiers to generate evaluation subset combinations, and partition identifiers are bound to each round of data partitioning conditions, thus forming multiple sets of data partitioning conditions. These data partitioning conditions characterize the fluctuations in model performance under different combinations of training and evaluation subsets. Stratified sampling can be represented by a case stratification indicator matrix and a sampling mask, and its expression is:

[0053]

[0054] in, Indicates the round-based index. This represents a subset of the assessment data pool grouped by cause of action identifier. Indicates a set of causes of action. Indicate the cause of action The proportion of samples allocated to the evaluation subset in this round. This represents the randomness control seed for that round. The sampling mask for this round indicates whether each sample enters the evaluation subset or the control subset. This process, by independently sampling each case subset in each round while maintaining a consistent proportion, makes the evaluation fluctuations across different rounds more reflective of small sample sensitivity rather than case proportion drift. Under each data partition condition, feature extraction is performed on the evaluation subset using a feature extraction framework to obtain feature extraction results. The feature extraction results are then aligned feature-wise with the normalized feature annotation results. Feature-wise alignment uses the consistency of normalized feature names and the text span overlap meeting a threshold as matching conditions. This allows for the statistical analysis of the hit and error rates for each normalized feature. The text span overlap can be measured using the intersection-union ratio (IUGR), whose expression is:

[0055]

[0056] in, This represents the actual text span range of the annotation. This represents the predicted text span range in the feature extraction results. This represents the length of the intersection of two intervals. This represents the union length of two intervals; this metric, by comparing the overlap between the predicted boundary and the true boundary, achieves tolerance for boundary jitter and ensures the objectivity of feature-by-feature alignment. When acquiring the performance statistics of each normalized feature under different data partitioning conditions and summarizing them to characterize performance stability, the server side calculates the precision, recall, and harmonic metric for each normalized feature under each round of data partitioning conditions, and summarizes the metric sequences from all rounds into a performance statistics set for that normalized feature. This performance statistics set reflects the performance fluctuation of the normalized feature under different data partitioning conditions; [The text then abruptly shifts to a different topic:] ...taking the first... Taking a wheel as an example, a certain normalized feature The harmonic index can be expressed as:

[0057]

[0058] in, Indicates the first Wheel-mounted normalized features The accuracy, derived from normalized features The number of correct predictions and normalized features The total predicted quantity is determined together. Indicates the first Wheel-mounted normalized features The recall rate, derived from normalized features The number of correct predictions and normalized features The total actual quantity is determined together. The numerical stability term is used to avoid the denominator being zero. The values ​​are positive and much smaller than conventional counting scales. After summarizing multiple rounds of indicators, performance stability does not use a single mean but rather lower-bound guided robust statistics to characterize the worst performance trend under extreme split conditions. For example, conditional risk value is used as performance stability, and its expression is:

[0059]

[0060] in, Indicates the total number of rounds. This refers to the tail proportion parameter, used to specify the percentage of rounds where the worst performance is highlighted. The smaller the value, the more attention is paid to the worst-case scenario where the value is within the range. Indicates will Sort by size from smallest to largest and take the first few. The index set consists of several rounds of indexes. This stability directly characterizes the lower bound of the performance of normalized features under small sample partitioning perturbations by averaging only the results of a few rounds with the worst tail, thus avoiding the mean being inflated by a few accidental high-scoring rounds and masking the true shortcomings. When obtaining the sample distribution information of each normalized feature in the training sample set and jointly analyzing it with the performance stability to form a resource-aware description, the server side counts the number of samples and cross-case coverage of each normalized feature in the training sample set. The number of samples refers to the total number of times the normalized feature appears in the training sample set, and the cross-case coverage refers to the proportion of cases covered by the normalized feature in the case set. To ensure that statistics of different magnitudes are comparable, the number of samples and cross-case coverage are mapped to sparsity and combined with the performance stability to form a resource-aware description. Sparsity can be expressed as:

[0061]

[0062] in, Representing normalized features The number of samples in the training sample set Representing normalized features The number of cases covered Indicates the size of the set of cause-of-cases. and The sparsity fusion weight is used to balance the contribution of scarce sample quantity and scarce cross-case coverage to the overall sparsity. The weight value is non-negative and normalizable. This sparsity avoids high-frequency features from masking long-tail features by using logarithmic compression on the sample quantity. At the same time, it emphasizes that features that appear only in a few cases are more likely to cause generalization instability by emphasizing cross-case coverage penalty. The resource-aware description takes the combination of performance stability and sparsity as input and explicitly expresses that normalized features with low performance lower bound and sparse sample distribution need to be enhanced first. When performing enhancement priority ranking and determining the features to be enhanced based on resource-aware description, the server first calculates an enhancement priority score for each normalized feature and generates an enhancement priority list from high to low based on the enhancement priority scores. Then, the normalized features located in a preset number of columns of the enhancement priority list are determined as the features to be enhanced. The enhancement priority list is used to determine the set of normalized features that should be prioritized for enhancement under limited enhancement resources. The enhancement priority score can adopt a multiplicative coupling of the degree of insufficiency of the lower bound of stability and the degree of sparse distribution to highlight the normalized features that simultaneously satisfy both types of shortcomings. Its expression is:

[0063]

[0064] in, Representing normalized features Performance stability Indicates the degree to which the lower bound of performance is insufficient. Representing normalized features sparsity, This represents a tradeoff index, used to adjust whether the bias is more towards insufficient performance lower bound or more towards sparse distribution. The higher the value, the more emphasis is placed on performance stability factors. This score avoids the dominance of extreme values ​​of a single factor in the ranking through multiplicative coupling, so that the enhancement priority is significantly increased only when both low performance stability and sparse sample distribution are true. This concentrates enhancement resources on the normalized features that are most likely to constitute the real shortcomings of the system. The preset number of columns is used to limit the scale of the number of features to be enhanced. The preset number of columns is determined by the engineering resource budget or the upper limit of the enhancement round. The final output is the set of features to be enhanced and their corresponding normalized feature names, case coverage information and performance stability information, so that subsequent simulation enhancement processing can be generated and screened in a targeted manner around the features to be enhanced and maintain consistency with the closed loop of the previous evaluation.

[0065] S150. For the features to be enhanced, while keeping the normalized feature system unchanged, use the simulation enhancement model to generate simulated case text, and perform adaptive evaluation and screening on the simulated case text based on the labeled sample set to obtain the simulation sample set.

[0066] Specifically, when constructing the simulation enhancement constraint set while maintaining the normalized feature names, semantic definitions, and applicable causes of action in the normalized feature system, the server generates a corresponding simulation enhancement constraint set for each feature to be enhanced and solidifies the simulation enhancement constraint set into a computable set of constraint items. The simulation enhancement constraint set includes at least semantic role constraints, context co-occurrence constraints, and cause of action adaptation constraints. Among them, semantic role constraints are used to limit the law enforcement semantic position of the feature to be enhanced in the case text, such as illegal acts, involved substances, quantity of operations, and location. Normalized features such as time are defined as roles such as action triggering, object reference, quantity modification, spatial positioning, and temporal positioning, respectively. Context co-occurrence constraints are used to limit the co-occurrence set, relative distance, and order of the feature to be enhanced with other normalized features, avoiding co-occurrence patterns in the generated text that conflict with real law enforcement narratives. Case type adaptation constraints are used to limit the feature to be enhanced to be generated only within the scope of its applicable case type and to appear in accordance with the common narrative template of that case type, thereby eliminating cross-case type context drift. To ensure that the constraint set can be used uniformly by the simulation enhancement model and subsequent screening, the constraint set is represented as a weighted sum of multiple penalties, and this weighted sum is used as the constraint cost of the candidate text. The expression for the constraint cost is:

[0067]

[0068] in, This indicates the candidate simulated case text. Indicates the feature identifier to be enhanced. Indicates the cause of action. The semantic role constraint cost is used to measure whether the semantic roles of the features to be enhanced in the candidate text are consistent with the preset semantic roles. The context co-occurrence constraint cost is used to measure whether the co-occurrence relationship between the feature to be enhanced and other normalized features in the candidate text falls within the allowable range. This represents the cost of case cause adaptation constraints, used to measure whether candidate texts satisfy the requirements of case cause narrative style and completeness of case cause elements. , , The weight parameters represent the costs of the three types of constraints. These weights are non-negative and determined by the statistical stability of the labeled sample set. This expression maps semantic roles, co-occurrence relationships, and cause-of-fact adaptation constraints to the same cost scale, ensuring consistent adjudication in subsequent generation and selection processes around the same constraint objective. When constructing a demonstration sample set based on the labeled sample set and driving the simulation enhancement model to generate candidate simulation case texts, the server first retrieves labeled samples from the labeled sample set that are consistent with the identifier of the feature to be enhanced and the applicable cause of action. Consistency of normalized feature names and cause-of-fact identifiers is a necessary condition for retrieval. From each labeled sample, a case text fragment containing the text span of the feature to be enhanced and its context fragments are extracted to construct a demonstration sample unit. This demonstration sample unit simultaneously retains the text span position of the feature to be enhanced, the set of normalized feature names co-occurring with the feature to be enhanced, and sentence / segment boundary information, thereby enabling the simulation enhancement model to... The model learns the local structural patterns in real law enforcement narratives; it applies repetition template suppression and coverage constraints to multiple demonstration sample units, ensuring that the demonstration sample set simultaneously covers different narrative sentence structures, different co-occurrence combinations, and different length ranges, avoiding excessive homogeneity in the demonstration sample set that leads to a single generation pattern; subsequently, the demonstration sample set and the simulation enhancement constraint set are jointly applied to the generation process of the simulation enhancement large model, maximizing stylistic consistency with the demonstration sample set while satisfying the constraint set. The generation objective can be expressed as a joint score of the constraint cost imposed on the basic generation probability and the demonstration consistency reward during sampling, and the expression for the joint score is:

[0069]

[0070] in, The word sequence representing the candidate simulated case text. Indicates the generation length. This indicates that the simulation enhancement model has parameters Next to the The conditional generation probability of each word element The feature to be enhanced is identified as And the cause of action is marked as The set of exemplary samples, The consistency reward is used to measure the similarity between candidate texts and the sample set in terms of wording style, sentence structure, and element combination. This indicates the reward intensity parameter. This indicates the aforementioned constraint cost; the joint scoring, by simultaneously encouraging real-world text and penalizing violations of constraints, ensures that candidate simulated case texts can both cover the features to be enhanced and maintain the stability and consistency of semantic roles and case context. When performing text normalization and feature localization on the candidate simulated case text to obtain the first candidate simulated case text, the server first applies character normalization rules consistent with those of the real case text, including unified writing of numbers and units, unified full-width and half-width characters, unified punctuation, and redundant whitespace removal, to ensure that subsequent feature localization is reproducible at the character-level boundaries and consistent with the encoding rules of the labeled sample set. Then, feature localization is performed to determine the text span position of the feature to be enhanced in the candidate simulated case text and to verify whether the text span satisfies semantic role constraints and context co-occurrence constraints. Feature localization does not rely on single string matching but uses a joint judgment of semantic consistency and fragment boundary interpretability between the candidate fragment and the normalized feature semantic prototype, enabling stable localization of synonymous rewriting. All possible spans in the candidate simulated case text are used to form a candidate span set, and the normalized feature matching probability and role consistency score are calculated for each span. The span that meets the threshold and has the highest score is selected as the localization span of the feature to be enhanced. The expression for selecting the localization span is:

[0071]

[0072] in, Indicates candidate simulation case text The candidate span set in This indicates the text span of the feature to be enhanced obtained from the localization. Indicates span Feature identifier to be enhanced The matching probability is calculated by the shared semantic encoder and the normalized feature semantic prototype set. This represents the semantic role consistency difference, used to measure the span. The degree to which the role in its context fits the preset role of the feature to be enhanced. This indicates that the compression function is used to map role consistency differences to intervals. Indicates the feature identifier to be enhanced The corresponding minimum matching probability threshold, This represents the context co-occurrence consistency discriminant. This indicates that the context co-occurrence constraint is satisfied; when a location span that meets the threshold cannot be found or the location span violates the semantic role constraint, the candidate simulated case text is directly removed, thereby obtaining the first candidate simulated case text and ensuring that it has a locationable feature text span to be enhanced and an interpretable semantic role. When constructing a true distribution baseline description based on the labeled sample set and selecting the second candidate simulated case text, the server first statistically forms a true distribution baseline description on the labeled sample set. This true distribution baseline description characterizes the statistical regularities of the real case text in terms of text length, normalized feature density, and semantic spatial location, ensuring that the selection of simulated samples relies on the empirical distribution of real data rather than fixed manual standards. Text length is characterized by character length or word length; normalized feature density is characterized by the number of occurrences of normalized features per unit length; and semantic similarity is characterized by the distance between the candidate text and the real text in the semantic vector space. For each first candidate simulated case text, the corresponding three types of characteristics are calculated and compared with the true distribution baseline description to determine whether it falls within the allowable interval of the true distribution. The allowable interval can be constrained by a robust quantile interval and Mahalanobis distance, thereby simultaneously controlling one-dimensional statistical deviation and multi-dimensional correlation deviation. The expression for the joint constraint is:

[0073]

[0074] in, Indicates whether to retain the first candidate simulated case text. The instructions result, Indicates text length characteristics. This represents the normalized feature density characteristic. and These represent the lower and upper quantiles of the actual text length distribution, respectively. and These represent the lower and upper quantiles of the true normalized feature density distribution, respectively. The semantic vector representation of the candidate text is obtained by aggregating the entire text using a shared semantic encoder. The mean vector representing the semantic vector of the real text. The covariance matrix representing the semantic vectors of the real text. Denotes the inverse matrix of the covariance matrix. Represents the semantic vector dimension. Describing the degrees of freedom as And the confidence level is The chi-square threshold is used to filter out candidate texts with abnormal length and density by using quantile intervals and to filter out candidate texts that deviate from the real text manifold in the semantic space by using Mahalanobis distance. This eliminates the target candidate simulated case texts that do not meet the real distribution constraints and obtains the second candidate simulated case text. When performing an adaptive multi-index comprehensive evaluation on the second candidate simulated case text and solidifying it into a simulated sample set, the server side constructs an evaluation index system with reference to the labeled sample set and adaptively determines the weights of each index. This ensures that the comprehensive evaluation guarantees both semantic consistency and feature coverage requirements, while avoiding subjective bias caused by manually fixed weights. The evaluation index includes at least semantic consistency, feature coverage, structural annotability, and diversity. The semantic consistency index measures the consistency between the second candidate simulated case text and the real case text in terms of semantic manifold. The feature coverage index measures the coverage strength and co-occurrence pattern coverage of the features to be enhanced in the simulated sample set. The structural annotability index measures the locatability stability of the text span boundaries. The diversity index suppresses templated repetition within the simulated sample set. The adaptive weights are determined by the stability and discriminability of each index in the labeled sample set. After normalizing each index, a weighted sum is obtained to obtain the comprehensive score. The expression for the comprehensive score is:

[0075]

[0076] in, Indicates the second candidate simulated case text The overall score, Indicates the number of indicators. Indicates the first Each indicator for text The original score, Indicates the first The median benchmark of each indicator on the labeled sample set Indicates the first The absolute median difference of each indicator on the labeled sample set is used to robustly characterize the degree of dispersion. Represents the numerically stable term. Indicates the first Adaptive weights for each indicator The weighted temperature parameter is used to control the sharpness of the weight distribution. Indicates the first The reliability score of each indicator is calculated by combining the stability of the indicator on the labeled sample set with its ability to distinguish the features to be enhanced. This comprehensive score is obtained by first performing robust standardization on the indicator scores to eliminate magnitude differences, and then adaptively allocating weights according to reliability, so that the indicators that are more stable and better reflect the quality of the features to be enhanced have a higher proportion. Finally, simulated case texts that meet the quantity budget and coverage constraints are selected from high to low according to the comprehensive score. Each selected text, along with the corresponding feature to be enhanced identifier, case identifier, and text span information obtained from the location, are solidified to form a simulated sample set. This ensures that the simulated sample set has a quality foundation that is localizable, interpretable, and alignable under the constraints of the normalized feature system, and can be used for subsequent joint training.

[0077] S160. A type divide-and-conquer and boundary arbitration mechanism is introduced between the simulation sample set and the labeled sample set to construct a joint training sample set. Under the constraints of a fixed model structure, fixed hyperparameter configuration, and deterministic training strategy, the simulation enhancement model is updated through the joint training sample set to output structured maritime illegal event characteristic results.

[0078] Specifically, when dividing normalized features into a set of features to be enhanced and a set of features not to be enhanced, and performing type-divide-and-conquer processing on the normalized feature annotations in the simulation sample set, the server side determines the set of features to be enhanced based on the enhancement priority list, and assigns the remaining normalized features to the set of features not to be enhanced, so that the subsequent supervision signal is controllable in the feature type dimension. The type-divide-and-conquer processing splits the normalized feature annotations in each simulated case text according to the set to which the normalized features belong. The simulation annotations belonging to the set of features to be enhanced are retained as target supervision information, while the simulation annotations belonging to the set of features not to be enhanced are removed to avoid the simulation generation error from causing negative transfer to stable features. The key to this processing is to keep the normalized feature system unchanged, that is, the normalized feature name, semantic definition and applicable case do not change. Type-divide-and-conquer only changes the scope of supervision for the simulation sample set to enter the training, thereby strictly focusing the role of simulation enhancement on the features to be enhanced at the training target level. When generating candidate backfill labels for the non-enhanced feature annotations removed from the simulation sample set, the server calls the feature extraction framework to generate a set of predicted labels for each simulated case text that has completed type divide-and-conquer. From this set, predicted labels whose normalized features belong to the non-enhanced feature set are selected as candidate backfill labels. Simultaneously, prediction confidence and text span position information are bound to the candidate backfill labels. The prediction confidence does not use simple Softmax output, but rather a joint characterization using temperature calibration and evidence distribution. This ensures that the confidence reflects both the relative advantage between categories and the magnitude of output uncertainty. The joint characterization expression is:

[0079]

[0080] in, This represents a simulated case text. Indicates the predicted text span. This represents the normalized feature category index corresponding to the predicted text span. This represents the total number of normalized feature categories, including non-feature categories. This represents the evidence parameter vector, where each dimension of the evidence parameter vector... The model represents the span Belongs to the The strength of evidence for a class Indicates the Dirichlet distribution. Let be a random variable representing the category probability vector; this expression interprets the model output as evidence rather than direct probabilities, thus deriving prediction confidence from the proportion of evidence, and when the total evidence... A smaller value implies higher overall uncertainty, thus providing a more stable basis for subsequent confidence grading; the evidence parameter vector is obtained by temperature calibration and nonlinear evidence mapping of the matching score, and the expression is:

[0081]

[0082] in, The feature extraction framework represents the span Belongs to the Unnormalized matching score of the class, Indicates the first The center calibration parameter for class scores is used to correct for differences in class scale. Indicates the first Temperature-like calibration parameters are used to adjust the sensitivity of the score-to-evidence ratio. This is used to map the calibrated scores to non-negative evidence. The preceding 1+ ensures that the parameter of each type of evidence is not less than 1 to satisfy the definition of the Dirichlet distribution parameter. The principle of this mapping is to transform the degree of advantage of the score relative to the center into evidence increment, so that the confidence is affected not only by the maximum class score, but also by the inter-class difference and the stability of the score scale. When performing boundary arbitration between candidate backfill annotations and feature annotations to be enhanced, the server side uses the same simulated case text as the scope, and retrieves all feature annotations to be enhanced for each candidate backfill annotation and judges whether there is a boundary conflict. Boundary arbitration emphasizes the priority of completely consistent conflicts. That is, when a candidate backfill annotation is completely consistent with any feature annotation to be enhanced in the text span position, the corresponding candidate backfill annotation is eliminated according to the priority principle of feature annotation to be enhanced, so as to ensure that the target supervision of the feature to be enhanced is not covered by the backfill supervision. In order to avoid the boundary judgment relying solely on strict equality, which leads to ineffectiveness to coding differences, the boundary arbitration simultaneously calculates the span equivalence judgment and character-level interval similarity, and determines the conflict type by using equivalence judgment as the main method and similarity as the auxiliary method. The span equivalence judgment and interval similarity are jointly expressed as:

[0083]

[0084] in, This indicates the text span range of the candidate backfill annotations. This represents the text span range for which features to be enhanced. and They represent The start and end character positions, and They represent The start and end character positions, Indicates an indicator function, This represents the intersection-union ratio (IU / U) of two text span intervals. This represents a similarity trigger threshold used to identify potential conflicts that are highly overlapping but of different types. This indicates the normalized feature name corresponding to the span. This represents the weight of similarity conflict items. The principle of this expression is to use strict boundary equality to ensure that completely consistent conflicts are forcibly arbitrated, while using highly overlapping but different types of conflict items to identify possible label competition risks, which facilitates subsequent statistical diagnosis and rule correction. However, the actual elimination action is only triggered for completely consistent conflicts, so as not to destroy the ability of overlapping text span and nested text span to coexist under different normalized feature names. When performing confidence grading on candidate backfill annotations retained through boundary arbitration based on predicted confidence, the server first constructs a confidence baseline distribution on the annotation sample set. Specifically, it performs inference on the annotation sample set and filters correctly predicted annotations to form a credible prediction set. Then, within the credible prediction set, it statistically analyzes the total evidence and mean confidence of the evidence parameter vector according to normalized feature categories. This allows for the determination of a certainty backfill threshold and a rejection threshold for each category. The threshold does not use fixed quantiles but rather a confidence lower bound guarantee. This requires that the certainty backfill annotations maintain a sufficiently high posterior lower bound at a given confidence level, thereby reducing threshold drift caused by small sample fluctuations. The posterior lower bound is determined based on the marginal Beta distribution corresponding to the Dirichlet distribution, expressed as:

[0085]

[0086] in, Indicate category The posterior lower bound, This represents the confidence level parameter, used to control the degree of conservatism of the lower bound. This represents the inverse cumulative distribution function of the Beta distribution. For predicting categories Evidence parameters, This is the sum of evidence parameters for the non-predictive categories; the principle behind this expression is to treat the prediction probability as a random variable and take its value at a confidence level. A conservative lower bound is established, meaning the lower bound will only increase when there is sufficient evidence and a clear class advantage; based on this lower bound, a classification is performed to ensure that the backfill annotation meets the requirements. Discard backfill markings to meet requirements The rest are fuzzy backfill labels and marked as gradient suppression locations, so that uncertain supervision does not drive parameter updates but can still retain its structural information for alignment and statistics. When merging samples from the labeled sample set with samples from the simulated sample set after type divide-and-conquer, boundary arbitration, and confidence leveling to construct a joint training sample set containing source mask information, the server directly writes the labeled sample set with the original labeling results into the joint training sample set, and writes the simulated sample set with the feature to be enhanced and the confidence backfilling label into the joint training sample set. At the same time, the text span position corresponding to the fuzzy backfilling label is written into the source mask field. The source mask field uses a multi-channel mask to simultaneously express information such as whether gradients are backpropagated, the strength of the supervision source, and the supervision type weight. The mask construction maps to the training unit space with the text span as the basic unit, so that subsequent loss calculation can accurately mask the fuzzy backfilling label. The principle of this merging is to explicitly encode the supervision source into the sample structure, so that the training stage can distinguish the reliability of different sources and perform differentiated gradient control without relying on external rules. This ensures that the joint training sample set both expands the effective supervision of the feature to be enhanced and does not introduce uncertain backfilling supervision into the parameter update path. Under the constraints of a fixed model structure, fixed hyperparameter configuration, and deterministic training strategy, when updating the simulation enhancement model based on the joint training sample set and outputting structured maritime illegal event feature results, the server side fixes the shared semantic encoder structure, span enumeration strategy, matching head structure, and all hyperparameter values. A deterministic training strategy is implemented by using a fixed randomness control seed and a fixed data iteration order, ensuring that the same joint training sample set corresponds to a unique and reproducible convergence trajectory. During training, a source-aware weighted loss and gradient blocking joint noise suppression supervision are used. The source-aware weighted loss maps labeled sample supervision, feature label supervision to be enhanced, confident backfill label supervision, and fuzzy backfill label supervision to different loss weights. An explicit gradient blocking operator is used at the fuzzy backfill label positions to prevent gradient backpropagation. The expression for the joint loss is:

[0087]

[0088] in, This indicates the parameters of the large simulation enhancement model. Represents the training batch sample set. The text indicating the case details, Represents the set of supervised annotations. Indicates text The constructed set of training units, where each training unit can correspond to a candidate text span or a span matching event. This indicates that the model supports the training units. The predicted output, Represents training unit The corresponding supervision objectives, Represents the cell-level loss function. This indicates a manually labeled supervised mask. This indicates the supervised mask for the annotation of features to be enhanced. This indicates that the backfill annotation supervision mask is confirmed. This indicates a mask for fuzzy backfill annotations. , , These represent the weighting coefficients for the three types of effective supervision, used to control the contribution of supervision from different sources to the overall gradient. The structural consistency coefficient represents the value of the fuzzy backfill annotation. This indicates a gradient blocking operator, which allows its forward values ​​to participate in the loss calculation but not... Gradient backpropagation is generated; the principle of this joint loss is to backpropagate gradients normally to reliable supervision positions to enhance the model's capabilities, while allowing fuzzy backfilled annotation positions to only constrain the forward shape of the predicted output without driving parameter updates, thereby maintaining the consistency of sample structure without introducing noisy gradients; after training, the updated simulation-enhanced large model is used to perform feature extraction inference on the case text to be processed, outputting structured maritime illegal event feature results consistent with the normalized feature system, and binding text identifiers, normalized feature names, start character positions, and end character positions to each output, so that the results can be directly entered into the structured processing link of subsequent law enforcement business processes.

[0089] This application also provides a simulation-enhanced feature extraction device for illegal events in small-sample maritime areas, referring to... Figure 2 , Figure 2This application provides a schematic diagram of a simulation-enhanced feature extraction device for small-sample maritime illegal events, which is a server. The server includes an acquisition module 21 and a processing module 22. The acquisition module 21 acquires case texts of various types of maritime illegal events and performs semantic merging and legal attribute verification on candidate features with business orientation in the case texts to construct a normalized feature system. The processing module 22 selects a set of labeled samples from each case text based on the normalized feature system and performs structural reconstruction and unified encoding on the labeled sample set to generate a training sample set. The processing module 22 further introduces a text-span-based feature extraction framework to extract features from the case texts based on the training sample set, obtaining feature extraction results containing multiple normalized features. The processing module 22 also uses resources to... The robustness assessment strategy of perception characterizes the performance stability of each normalized feature under different data partitioning conditions, and determines the enhancement priority of each normalized feature by combining the sample distribution information of each normalized feature in the training sample set, so as to determine the feature to be enhanced according to the enhancement priority; the processing module 22 is also used to generate simulated case text using the simulation enhancement model for the feature to be enhanced, while keeping the normalized feature system unchanged, and perform adaptive evaluation and screening on the simulated case text based on the labeled sample set to obtain the simulation sample set; the processing module 22 is also used to introduce a type divide-and-conquer and boundary arbitration mechanism between the simulation sample set and the labeled sample set to construct a joint training sample set, and update the simulation enhancement model through the joint training sample set under the constraints of fixed model structure, fixed hyperparameter configuration and deterministic training strategy, so as to output the structured maritime illegal event feature results.

[0090] This application also provides an electronic device, with reference to... Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: at least one processor 31, at least one network interface 34, a user interface 33, a memory 35, and at least one communication bus 32.

[0091] The communication bus 32 is used to enable communication between these components.

[0092] The user interface 33 may include a display screen and a camera. Optionally, the user interface 33 may also include a standard wired interface and a wireless interface.

[0093] The network interface 34 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0094] The processor 31 may include one or more processing cores. The processor 31 connects to various parts of the server via various interfaces and lines, executing instructions, programs, code sets, or instruction sets stored in the memory 35, and calling data stored in the memory 35 to perform various server functions and process data. Optionally, the processor 31 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 31 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 31 and may be implemented as a separate chip.

[0095] The memory 35 may include random access memory (RAM) or read-only memory. Optionally, the memory 35 may include a non-transitory computer-readable storage medium. The memory 35 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 35 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 35 may also be at least one storage device located remotely from the aforementioned processor 31. Figure 3 As shown, the memory 35, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a simulation-enhanced method for extracting features of illegal events in small-sample marine areas.

[0096] exist Figure 3In the electronic device shown, the user interface 33 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 31 can be used to call an application stored in the memory 35 for a simulation-enhanced method for extracting features of illegal events in a small number of sea areas. When executed by one or more processors, the electronic device executes one or more methods as described in the above embodiments.

[0097] This application also provides a non-transitory computer-readable storage medium storing instructions. When executed by one or more processors, these instructions cause an electronic device to perform one or more of the methods described in the above embodiments.

[0098] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for extracting features of illegal events in small-sample maritime areas based on simulation enhancement, characterized in that, The method includes: The case texts of various types of maritime illegal cases are obtained, and the candidate features with business orientation in the case texts are semantically merged and legal attribute verified in order to construct a normalized feature system. Based on the normalized feature system, a set of labeled samples is selected from each case text, and structural reconstruction and unified encoding are performed on the set of labeled samples to generate a set of training samples. Based on the training sample set, a feature extraction framework based on text span is introduced to extract features from the case text, resulting in feature extraction results containing multiple normalized features. Based on the feature extraction results, a resource-aware robustness evaluation strategy is introduced to characterize the performance stability of each normalized feature under different data partitioning conditions. Combined with the sample distribution information of each normalized feature in the training sample set, the enhancement priority of each normalized feature is determined, so as to determine the feature to be enhanced according to the enhancement priority. While keeping the normalized feature system and the training sample set unchanged, an evaluation data pool is constructed, and each sample in the evaluation data pool is bound with a sample identifier, a case identifier, and a normalized feature annotation result; Multiple rounds of data partitioning are performed on the evaluation data pool. Multiple sets of data partitioning conditions are generated in the case cause dimension. Under each data partitioning condition, the feature extraction framework is used to perform feature extraction processing on the corresponding evaluation subset. The obtained feature extraction results are aligned with the real annotation results feature by feature to obtain the performance statistics of each normalized feature under different data partitioning conditions. Summarize the performance statistics of each normalized feature under multiple data partitioning conditions, and characterize the performance stability of the normalized feature to changes in data partitioning based on the performance statistics. Obtain the sample distribution information of each normalized feature in the training sample set, and perform joint analysis of the sample distribution information and the performance stability to form a resource-aware description. The sample distribution information includes the number of samples and cross-case coverage. Based on the resource-aware description, the normalized features are sorted by enhancement priority. An enhancement priority list is generated according to the comprehensive order of performance stability and sample distribution sparsity. The normalized features located in the preset number of columns of the enhancement priority list are determined as the features to be enhanced. For the features to be enhanced, while keeping the normalized feature system unchanged, a simulation enhancement model is used to generate a simulation case text, and an adaptive evaluation and screening is performed on the simulation case text based on the labeled sample set to obtain a simulation sample set. A type-divide-and-conquer and boundary arbitration mechanism is introduced between the simulation sample set and the labeled sample set to construct a joint training sample set. Under the constraints of a fixed model structure, fixed hyperparameter configuration, and deterministic training strategy, the simulation enhancement model is updated through the joint training sample set to output structured maritime illegal event feature results.

2. The method for extracting features of illegal events in small-sample maritime areas based on simulation enhancement according to claim 1, characterized in that, The process of acquiring case texts of various types of maritime illegal activities and semantically merging and legally verifying the business-oriented candidate features in the case texts to construct a normalized feature system specifically includes: A case set consisting of multiple types of maritime illegal cases is determined, and a unique case identifier is bound to each case in the case set. Based on the case identifier, case texts consistent with the case identifier are obtained from the case management system, law enforcement document system and case entry system. Each case text is bound to a text identifier, case identifier, generation time identifier and source identifier to form a case text library. The case texts in the case text database are subjected to character normalization, noise fragment removal, and sentence segmentation to generate normalized case texts bound with sentence segment identifiers and location identifiers; Based on the standardized case text, word segmentation, part-of-speech tagging, and stop word filtering are performed on the sentence and segment sequences. Statistical significance analysis is then performed on the processed word units for different case types to form a preliminary set of candidate features. The initial screening candidate feature set is subjected to synonym merging, entity morphology merging, and cross-case co-occurrence aggregation to generate a candidate feature context profile containing a co-occurrence word set, a dependency relationship fragment set, and a trigger phrase set. Under the constraints of a pre-established top-level feature set, semantic discrimination processing is performed on each candidate feature in conjunction with the candidate feature context profile to determine the mapping relationship between the candidate features and the top-level features, and a draft of a normalized feature system covering multiple causes of action is formed accordingly. Based on the set of legal elements, the normalized features in the draft normalized feature system are subjected to legal attribute verification processing. By verifying the consistency between the normalized features and legal elements and applicable causes of action, the mapping relationship is corrected and the conflicts of synonyms and the redundancy of synonyms are eliminated, so as to form a normalized feature system that includes feature identifiers, feature definitions, a set of applicable causes of action, a set of trigger phrases and a set of negative example constraints.

3. The method for extracting features of illegal events in small-sample maritime areas based on simulation enhancement according to claim 1, characterized in that, Based on the normalized feature system, a set of labeled samples is selected from each case text, and structural reconstruction and unified encoding are performed on the labeled sample set to generate a training sample set, specifically including: The case texts are grouped by cause of action, and within each cause of action group, the frequency of occurrence and the number of sentences covered by each normalized feature in the normalized feature system are statistically analyzed to form a feature distribution profile of the cause of action dimension. Under the constraint of the feature distribution profile, case texts containing multiple normalized features and semantically complete are selected from each case group as candidate labeled texts, and the number of candidate labeled texts in each case group is kept consistent, so as to construct a balanced labeled sample set across case groups. Based on the normalized feature system, the labeled sample set is subjected to fine-grained labeling. The normalized feature name is used as the unique label. The text fragments in the case text that are semantically consistent with the normalized feature name are labeled at the span level. Each labeling result is bound to a text identifier, a normalized feature name, a start character position, and an end character position. The labeled samples are restructured to form a flattened sample structure containing text identifiers, case identifiers, text content fields, and a feature list field consisting of multiple feature objects. The feature objects include normalized feature names, start character positions, and end character positions. Based on the flattened sample structure, a unified encoding process is performed on the labeled sample set to generate a globally unique sample identifier for each labeled sample, which includes the case type identifier and the original text identifier. Dictionary mapping is performed on the normalized feature names to ensure the encoding consistency of the same normalized feature in different samples, while verifying the consistency between the feature span position and the text content. After completing the unified coding and consistency verification, the labeled sample set is divided into training sample set, verification sample set and test sample set according to the case type identifier.

4. The method for extracting features of illegal events in small-sample maritime areas based on simulation enhancement according to claim 3, characterized in that, Based on the training sample set, a text span-based feature extraction framework is introduced to extract features from the case text, resulting in feature extraction results containing multiple normalized features, specifically including: By using a shared semantic encoder to uniformly encode the case text and the normalized feature name respectively, the context representation sequence of the case text and the set of normalized feature semantic prototypes are obtained. In the context representation sequence, candidate text spans are enumerated based on a preset span length, and the candidate text spans are subjected to legality constraint filtering to form a candidate text span set; For each candidate text span in the candidate text span set, a corresponding text span semantic representation is constructed from the context representation sequence. The text span semantic representation includes semantic information within the span, span start boundary information, span end boundary information, and span length information. The semantic representation of the text span is semantically matched with the set of normalized feature semantic prototypes to obtain the matching result of the candidate text span on the full set of normalized features, and the normalized feature name or non-feature category corresponding to the candidate text span is determined based on the matching result. Aggregation and conflict resolution processes are performed on candidate text spans that are determined to be normalized features. While preserving overlapping and nested text spans, feature extraction results containing multiple normalized features are generated. Each normalized feature in the feature extraction results is bound to a text identifier, normalized feature name, and text span position information consistent with the training sample set.

5. The method for extracting features of illegal events in small-sample maritime areas based on simulation enhancement according to claim 1, characterized in that, For the features to be enhanced, while keeping the normalized feature system unchanged, a simulated case text is generated using a large-scale simulation enhancement model. Adaptive evaluation and filtering are then performed on the simulated case text based on the labeled sample set to obtain a simulated sample set. Specifically, this includes: While keeping the normalized feature names, semantic definitions and applicable causes of action unchanged in the normalized feature system, a set of simulated enhancement constraints is constructed for each feature to be enhanced. The set of simulated enhancement constraints is used to limit the semantic role, context co-occurrence relationship and cause of action applicable range of the feature to be enhanced in the case text. Based on the labeled samples in the labeled sample set that are consistent with the features to be enhanced and the applicable cause of action, the corresponding case text and context fragments are extracted to construct a demonstration sample set. The demonstration sample set and the simulation enhancement constraint set are used together as the input of the simulation enhancement large model to generate the candidate simulation case text containing the features to be enhanced. The candidate simulated case texts are subjected to text normalization and feature localization processing. Candidate simulated case texts that cannot locate the span of the text to be enhanced or do not meet the semantic role constraints are eliminated to obtain the first candidate simulated case text. Based on the labeled sample set, a true distribution benchmark description is constructed, and the text length characteristics, normalized feature density characteristics and semantic similarity characteristics of the first candidate simulated case text are compared with the true distribution benchmark description. Target candidate simulated case texts that do not meet the true distribution constraints are eliminated to obtain the second candidate simulated case text. An adaptive multi-index comprehensive evaluation process is performed on the second candidate simulated case text to determine the simulated case text that meets the requirements of semantic consistency and feature coverage. The simulated case text, along with the corresponding feature identifier to be enhanced, case identifier, and text span information, are then solidified to form the simulated sample set.

6. The method for extracting features of illegal events in small-sample maritime areas based on simulation enhancement according to claim 1, characterized in that, The method involves introducing a type-divide-and-conquer and boundary arbitration mechanism between the simulated sample set and the labeled sample set to construct a joint training sample set. Under constraints of a fixed model structure, fixed hyperparameter configuration, and a deterministic training strategy, the simulation-enhanced large model is updated using this joint training sample set to output structured maritime illegal event characteristic results. Specifically, this includes: The normalized features are divided into a set of features to be enhanced and a set of features not to be enhanced. While keeping the normalized feature system unchanged, the normalized feature annotations in the simulation sample set are subjected to type divide-and-conquer processing, and the simulation annotations belonging to the set of features to be enhanced are retained as target supervision information. For the non-features to be enhanced that have been removed from the simulation sample set, the feature extraction framework is used to generate predicted labels for the corresponding simulation case text, and the predicted labels whose normalized features belong to the non-features to be enhanced set are used as candidate backfill labels. Boundary arbitration is performed between the candidate backfill annotation and the feature annotation to be enhanced. When the candidate backfill annotation and any feature annotation to be enhanced are completely consistent in the text span position, the corresponding candidate backfill annotation is eliminated according to the priority principle of the feature annotation to be enhanced. Candidate backfill labels retained after boundary arbitration are processed according to prediction confidence, and divided into sure backfill labels, fuzzy backfill labels and discarded backfill labels. Among them, sure backfill labels are used as effective supervision in training, discarded backfill labels are removed, and fuzzy backfill labels are marked as gradient suppression positions. The samples in the labeled sample set are merged with the samples in the simulation sample set after type divide-and-conquer, boundary arbitration and confidence level processing to construct a joint training sample set containing source mask information. Under the constraints of a fixed model structure, fixed hyperparameter configuration, and deterministic training strategy, the simulation enhancement model is updated based on the joint training sample set. Noise supervision is suppressed by blocking gradient backpropagation at fuzzy backfill annotation positions and ensuring normal backpropagation of gradients at the annotation positions of features to be enhanced and the confirmed backfill annotation positions, so as to output the structured marine illegal event feature results.

7. A device for extracting features of illegal events in small-sample maritime areas based on simulation enhancement, characterized in that, The apparatus is used to perform the simulation-enhanced feature extraction method for illegal maritime incidents based on any one of claims 1 to 6, wherein the apparatus includes an acquisition module and a processing module, wherein... The acquisition module is used to acquire case texts of various types of maritime illegal cases, and to perform semantic merging and legal attribute verification on the business-oriented candidate features in the case texts in order to construct a normalized feature system. The processing module is used to select a set of labeled samples from each case text based on the normalized feature system, and perform structural reconstruction and unified encoding processing on the set of labeled samples to generate a set of training samples. The processing module is also used to introduce a text span-based feature extraction framework on the basis of the training sample set to extract features from the case text and obtain feature extraction results containing multiple normalized features. The processing module is further configured to characterize the performance stability of each normalized feature under different data partitioning conditions by introducing a resource-aware robustness evaluation strategy based on the feature extraction results, and determine the enhancement priority of each normalized feature by combining the sample distribution information of each normalized feature in the training sample set, so as to determine the feature to be enhanced based on the enhancement priority. The processing module is further configured to generate simulated case text using a large simulation enhancement model for the features to be enhanced, while keeping the normalized feature system unchanged, and perform adaptive evaluation and filtering on the simulated case text based on the labeled sample set to obtain a simulated sample set. The processing module is further configured to introduce a type divide-and-conquer and boundary arbitration mechanism between the simulation sample set and the labeled sample set to construct a joint training sample set, and update the simulation enhancement model through the joint training sample set under the constraints of a fixed model structure, fixed hyperparameter configuration and deterministic training strategy, so as to output structured maritime illegal event feature results.

8. An electronic device, characterized in that, The electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Legal event detection model construction method based on double prototype sources and application

    CN118445621A

  • Natural language to low code conversion method based on multi-modal reinforcement learning

    CN121257475A