Method and device for training context sensing model

Through multi-level feature processing and feature fusion technology, the context-aware model is trained, which solves the problems of poor adaptability to targets at different scales and insufficient semantic information richness, and achieves stronger feature expression and semantic understanding capabilities.

CN120046598APending Publication Date: 2025-05-27XIAN SECLOVER INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510061206.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing context-aware model is not very adaptable to targets at different scales, has weak feature expression ability, and is difficult to effectively combine local and global information, and the richness of semantic information is insufficient.

Method used

By setting feature keywords and feature levels, multi-level feature capture, verification, feature hierarchy fusion and feature correlation reconstruction are carried out on the original text sample to form feature data with global correlation, and input them into the initial context-aware model for training to obtain the target context-aware model.

Benefits of technology

It enhances the adaptability of the model to targets at different scales, effectively combines local and global information, improves the richness of semantic information, and improves the feature expression ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046598A_ABST
    Figure CN120046598A_ABST
Patent Text Reader

Abstract

The invention discloses a context sensing model training method and device, and the method comprises the steps: carrying out the multi-level feature capture of an original text sample according to a feature keyword and a feature level number, and obtaining first feature data; verifying the first feature data according to a preset verification rule, and taking the first feature data passing the verification as second feature data; performing feature level fusion on the second feature data to obtain third feature data; performing feature correlation reconstruction on the third feature data to obtain fourth feature data with global correlation; and inputting the fourth feature data into a preset initial context sensing model for model training to obtain a target context sensing model. According to the scheme, the features of different levels can be effectively fused, the adaptability of the model to targets of different scales is enhanced, local and global information is effectively combined, the feature pairs are input and processed, the expression ability of the feature pairs is enhanced, and the richness of semantic information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular, to a method and device for training a context-aware model. Background Art

[0002] In a typical interaction scenario, when a user interacts with a large language model, the large language model uses an enhanced context-aware model to enhance the interaction information. The enhanced information helps the large language model understand and parse the user's query information, comprehensively improving the interaction effect and physical experience.

[0003] Context-aware technology has demonstrated remarkable efficiency and capabilities in current application scenarios. Because this technology can sense and understand the user's current environment, status, and tasks, providing personalized services. For example, in smartphone applications, context-aware computing can automatically provide nearby restaurant recommendations based on the user's location changes or adjust the screen brightness to protect eyesight. In addition, context-aware technology is also crucial in recommendation systems, enabling more accurate and personalized recommendations by combining the user's context information (such as time, location, device status). These applications not only enhance the intelligence and personalization level of the user experience but also greatly strengthen the interactivity between the user and the device. Therefore, context-aware technology plays an important role in improving service accuracy and user satisfaction and is an indispensable part of modern intelligent systems.

[0004] Current context-aware technologies mainly include intelligent processing and pattern recognition, context-aware computing, context-aware recommendation systems, etc. The advantage of these technologies is that they can provide more personalized and accurate services, making devices and applications more intelligent and responsive.

[0005] However, the context-aware function of the context-aware model has some drawbacks. For example, the model has poor adaptability to different scale targets, weak feature expression ability, inability to effectively combine local and global information, and insufficient richness of semantic information. Summary of the Invention

[0006] The present invention aims to at least solve the technical problems existing in the prior art. To this end, a first aspect of the present invention proposes a method for training a context-aware model, the method comprising:

[0007] Setting feature keywords and the number of feature levels according to the domain features of the original text sample;

[0008] Performing multi-level feature capture on the original text sample according to the feature keywords and the number of feature levels to obtain first feature data;

[0009] Verify the first feature data according to the preset verification rules, and use the first feature data that passes the verification as the second feature data;

[0010] Perform feature-level fusion on the second feature data to obtain third feature data;

[0011] Perform feature correlation reconstruction on the third feature data to obtain fourth feature data with global correlation;

[0012] Input the fourth feature data into a preset initial context-aware model for model training to obtain a target context-aware model.

[0013] Optionally, the step of capturing multi-level features from the original text sample according to the feature keywords and the number of feature levels to obtain the first feature data includes:

[0014] Capture local features from the original text sample through convolution operations according to the feature keywords and the number of feature levels, and capture global features through average pooling;

[0015] Combine the local features and the global features to obtain the first feature data.

[0016] Optionally, the step of performing feature-level fusion on the second feature data includes:

[0017] Aggregate the local information of the second feature data using depth convolution, capture the multi-scale context information of the second feature data using multi-branch depth convolution, and simulate the relationship between different channels in the second feature data using 1×1 pointwise convolution to obtain first relationship data;

[0018] Use the first relationship data as the weight of convolutional attention;

[0019] Extract low-level features and high-level features from the second feature data;

[0020] Use the weight to perform cross-layer fusion on the low-level features and the high-level features to obtain cross-layer fusion features.

[0021] Optionally, the step of performing feature correlation reconstruction on the third feature data to obtain fourth feature data with global correlation includes:

[0022] Obtain support-query feature pairs from the third feature data;

[0023] Obtain the context semantic information of the third feature data through a self-correlation reconstruction module, and capture local and global correlations for the support-query feature pairs in each semantic level;

[0024] Performing feature correlation reconstruction on the third feature data by using the local and global correlations to obtain fourth feature data with global correlation.

[0025] Optionally, the original text sample is data in the field of data security, and setting feature keywords and the number of feature levels according to the domain features of the original text sample includes:

[0026] Setting the number of feature levels of data security risks to three levels, where the first-level features include: data carrier risk, transmission risk, storage risk, application risk, sharing risk, destruction risk, etc.

[0027] Setting multiple second-level features for each of the first-level features; the second-level features of the storage risk include: physical storage risk, software storage risk.

[0028] Setting multiple third-level features for each of the second-level features; the third-level features of the software storage risk include software vulnerability risk, software version risk, software access risk, software original storage risk.

[0029] A second aspect of the present invention proposes a training device for a context-aware model, and the device includes:

[0030] A feature setting module, configured to set feature keywords and the number of feature levels according to the domain features of the original text sample;

[0031] A feature capture module, configured to perform multi-level feature capture on the original text sample according to the feature keywords and the number of feature levels to obtain first feature data;

[0032] A verification module, configured to verify the first feature data according to a preset verification rule and use the first feature data that passes the verification as second feature data;

[0033] A fusion module, configured to perform feature-level fusion on the second feature data to obtain third feature data;

[0034] A reconstruction module, configured to perform feature correlation reconstruction on the third feature data to obtain fourth feature data with global correlation;

[0035] A training module, configured to input the fourth feature data into a preset initial context-aware model for model training to obtain a target context-aware model.

[0036] Optionally, the feature capture module is specifically configured to:

[0037] Performing local feature capture on the original text sample by convolution operation according to the feature keywords and the number of feature levels, and capturing global features by average pooling.

[0038] Combine the local feature and the global feature to obtain first feature data.

[0039] Optionally, the fusion module is specifically configured to:

[0040] Aggregate the local information of the second feature data by using depth convolution, capture the multi-scale context information of the second feature data by using multi-branch depth convolution, and simulate the relationship between different channels in the second feature data by using 1×1 pointwise convolution to obtain first relationship data;

[0041] Use the first relationship data as the weight of convolutional attention;

[0042] Extract low-level features and high-level features from the second feature data;

[0043] Use the weight to perform cross-layer fusion on the low-level features and the high-level features to obtain cross-layer fusion features.

[0044] The third aspect of the present invention provides an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the training method of the context awareness model as described in the first aspect.

[0045] The fourth aspect of the present invention provides a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the training method of the context awareness model as described in the first aspect.

[0046] The embodiments of the present invention have the following beneficial effects:

[0047] In an embodiment of the present invention, feature keywords and the number of feature levels are set according to the domain characteristics of the original text sample; multi-level feature capture is performed on the original text sample according to the feature keywords and the number of feature levels to obtain first feature data; the first feature data is verified according to a preset verification rule, and the first feature data that passes the verification is used as second feature data; feature level fusion is performed on the second feature data to obtain third feature data; feature correlation reconstruction is performed on the third feature data to obtain fourth feature data with global correlation; the fourth feature data is input into a preset initial context-aware model for model training to obtain a target context-aware model. Through the multi-scale channel attention mechanism and the attention-induced cross-level fusion module, this solution can effectively fuse features at different levels and enhance the adaptability of the model to targets at different scales. Moreover, the dual-branch global context module works through two branches. One branch captures local features through convolution operations, and the other branch captures global features through average pooling, effectively combining local and global information. The feature enhancement module is used to enhance the feature expression of feature pairs. By inputting and processing the feature pairs, the expression ability of the feature pairs is enhanced. The self-correlation reconstruction module is used to capture the context semantic information of feature pairs and capture local and global correlations for each semantic level, improving the richness of semantic information. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of the steps of a method for training a context-aware model provided by an embodiment of the present invention;

[0049] Figure 2 A schematic diagram of a model construction process provided by an embodiment of the present invention;

[0050] Figure 3 is a block diagram of the structure of a training device for a context-aware model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "a plurality of" is two or more. Additionally, the use of "based on" or "according to" is meant to be open and inclusive, because a process, step, calculation, or other action "based on" or "according to" one or more of the stated conditions or values may, in practice, be based on additional conditions or values beyond those stated.

[0053] Figure 1 is a flowchart of the steps of a method for training a context-aware model provided by an embodiment of the present invention.

[0054] As Figure 1 shown, the method includes the following steps:

[0055] Step 101, set feature keywords and the number of feature levels according to the domain features of the original text sample.

[0056] In the process of feature extraction and use, features are often multi-level. For example, the first-level features are object form, object appearance, etc., and the next-level features of object shape include planar shape, spatial shape, etc.

[0057] Set multi-level features according to the domain features of the original text sample, including the keywords of the features, the number of feature levels, etc.

[0058] As an optional embodiment, if the original text sample belongs to the field of data security, set the number of feature levels of data security risks to three levels. Among them, the first-level features include: data carrier risk, transmission risk, storage risk, application risk, sharing risk, destruction risk, etc.

[0059] Set multiple second-level features for each of the first-level features; the second-level features of the storage risk include: physical storage risk, software storage risk;

[0060] Set multiple third-level features for each of the second-level features; the third-level features of the software storage risk include software vulnerability risk, software version risk, software access risk, software original storage risk.

[0061] In addition, for each feature, set the weight of the feature.

[0062] Step 102, perform multi-level feature capture on the original text sample according to the feature keywords and the number of feature levels to obtain first feature data.

[0063] In practical applications, data preprocessing is the first step in realizing context-based text information detection and extraction, including operations such as cleaning, tokenization, and stop word removal, to prepare for the subsequent processing of large language models (LLMs).

[0064] Specifically, common data integration and preprocessing techniques can be used to integrate and preprocess the dataset. Among them, data integration techniques include data processing techniques such as data association and data fusion. Preprocessing techniques include related techniques such as data standardization, formatting, and normalization.

[0065] Data preprocessing is the first step in realizing context-based text information detection and extraction, including operations such as data cleaning, data transformation, and data normalization to improve the performance and accuracy of the model.

[0066] Then, multi-level feature capture is performed on the original text samples according to the feature keywords and the number of feature levels to obtain the first feature data.

[0067] Specifically, the powerful semantic understanding ability of large language models can be utilized to deeply analyze the training data, extract key context features such as semantics, sentiment, and theme, to obtain the first feature data, providing a basis for subsequent analysis and processing. In this step, generally a privatized small and medium-sized large language model can be used, which can often achieve the expected effect while reducing costs.

[0068] As an optional embodiment, the multi-level feature capture of the original text samples according to the feature keywords and the number of feature levels to obtain the first feature data includes:

[0069] Step 1021: Capture local features of the original text samples through convolution operations according to the feature keywords and the number of feature levels, and capture global features through average pooling;

[0070] Step 1022: Combine the local features and the global features to obtain the first feature data.

[0071] In steps 1021 - 1022, there are various multi-level feature capture techniques. The dual-branch global context module technique can be used to perform multi-level feature capture. The dual-branch global context module uses rich global context information to enhance cross-layer fusion features. It works through two branches. One branch captures local features through convolution operations, and the other branch captures global features through average pooling. Finally, the local features and the global features are combined to obtain the first feature data.

[0072] The dual-branch global context module can enhance the context information of features by capturing local and global features, improving the model's recognition ability for large objects with global distribution and local context.

[0073] Step 103: Verify the first feature data according to the preset verification rules, and use the verified first feature data as the second feature data.

[0074] Pre-set the extraction rules for context features and the verification rules for the extraction rules. The extraction rules are to store the above multi-level features and their corresponding weights in a certain form and set them in a form that can be understood by the subsequent large language model.

[0075] The verification rules can specifically include: whether the form of feature extraction is correct, whether the feature format is correct, whether the extracted semantics deviate significantly, etc.

[0076] When verifying the first feature data, the original rule matching form can be used or the large language model can be used for verification. When using the large language model for verification, the verification rules can be incorporated into its input in the form of setting prompts, and the first feature data can be verified using the verification rules. The verified first feature data is used as the second feature data for subsequent processing, and the first feature data that fails the verification is discarded.

[0077] Step 104: Perform feature-level fusion on the second feature data to obtain the third feature data.

[0078] Specifically, through the multi-scale channel attention mechanism and the attention-induced cross-level fusion module, the effective fusion of features at different levels can be realized, enhancing the model's adaptability to targets at different scales.

[0079] The second feature data after feature extraction by the large language model may still be in the form of natural language. For example, the form of the feature extraction result of the large language model can be in a format like "xxx information belongs to xxx feature". The third feature data after performing feature-level fusion already has a standardized multi-level structure, which conforms to the above feature-level specification compared to when no feature extraction is performed. After feature-level fusion, the data format is more standardized, which is beneficial for subsequent feature correlation reconstruction.

[0080] As an optional embodiment, the performing feature-level fusion on the second feature data includes:

[0081] Step 1041: Aggregate the local information of the second feature data using depth convolution, capture the multi-scale context information of the second feature data using multi-branch depth convolution, and simulate the relationship between different channels in the second feature data using 1×1 pointwise convolution to obtain the first relationship data;

[0082] Step 1042: Use the first relational data as the weight of convolutional attention;

[0083] Step 1043: Extract low-level features and high-level features from the second feature data;

[0084] Step 1044: Use the weight to perform cross-layer fusion on the low-level features and the high-level features to obtain cross-layer fusion features.

[0085] In Steps 1041 - 1044, there are multiple multi-level feature fusion techniques, such as the cross-level fusion module technique induced by attention. The cross-level fusion module technique induced by attention fuses features at different levels by introducing features and a multi-scale channel attention module, enhancing the model's adaptability to targets at different scales. This module helps integrate multi-level features and emphasizes large objects in the global distribution as well as local contexts, avoiding small objects from being ignored.

[0086] The feature and multi-scale channel attention module consists of three parts: depth convolution aggregates local information, multi-branch depth convolution captures multi-scale context information, and 1×1 pointwise convolution simulates the relationship between different channels in the features. The output of the 1×1 pointwise convolution is used as the weight of convolutional attention to re-weight the input of the MSCA. This mechanism can effectively fuse cross-level features and use multi-scale information to mitigate the impact of scale changes.

[0087] The output of the feature and multi-scale channel attention module is the first relational data, which can be used as the weight of convolutional attention.

[0088] Since low-level features with higher spatial resolution require more computational resources than high-level features but contribute less to the performance of the deep integration model. Therefore, the multi-scale channel attention mechanism is only executed in high-level features.

[0089] Specifically, we use fi (i = 1, 2, 3) as low-level features and fi (i = 3, 4, 5) as high-level features. The cross-level fusion process can be described as follows:

[0090]

[0091] Among them, F ab represents the cross-layer fusion feature, M represents the weight of convolutional attention, F a represents the low-level feature, and F b represents the high-level feature.

[0092] For the high-level feature F b , first upsample it twice to the same scale as the low-level feature F a , and then combine it with Fa Add them up. After being processed by the feature and multi-scale channel attention module, the obtained feature is multiplied by F a Take the inverse after multiplication and add it to the upsampled F b Multiply them. Add the results of the two products obtained, then pass through a 3×3 convolutional layer, and then perform batch normalization and ReLU activation function. Finally, obtain the cross-layer fusion feature F ab .

[0093] Step 105: Perform feature correlation reconstruction on the third feature data to obtain fourth feature data with global correlation.

[0094] There are many feature correlation reconstruction techniques, such as self-correlation reconstruction module technology. The self-correlation reconstruction module technology is used to capture the contextual semantic information of feature pairs and capture local and global correlations for each semantic level, enhancing the semantic understanding ability of the model.

[0095] The self-correlation reconstruction module technology captures contextual semantic information through a new self-similarity module and establishes local and global correlations between support and query features at each semantic level.

[0096] Without feature correlation construction, the features have a hierarchy but lack global correlation, which is not conducive to the perception of subsequent contextual content with global correlation. However, the fourth feature data after feature correlation reconstruction has global feature correlation rather than local feature correlation.

[0097] As an optional embodiment, the performing feature correlation reconstruction on the third feature data to obtain fourth feature data with global correlation includes:

[0098] Step 1051: Obtain support-query feature pairs from the third feature data;

[0099] Step 1052: Obtain the contextual semantic information of the third feature data through the self-correlation reconstruction module and capture local and global correlations for the support-query feature pairs in each semantic level;

[0100] Step 1053: Use the local and global correlations to perform feature correlation reconstruction on the third feature data to obtain fourth feature data with global correlation.

[0101] In steps 1051 - 1053, the Query feature can be understood as the "query" feature, which is the feature vector for which we want to find matching or relevant information. In many applications, such as image retrieval, face recognition, or recommendation systems, the query feature represents the query conditions input by the user, or the feature representation of the target object for which we want to find similar items. The Support feature, on the other hand, can be understood as the "support" feature, which is a set of feature vectors used to compare with the query feature. During the training phase, the support features usually come from the training dataset and are used to help the model learn how to identify and match the query features. During the inference or application phase, the support features may come from a database or data pool and are used to compare with the query features to find the most similar items.

[0102] During execution, obtain the support - query feature pairs from the third feature data; obtain the context semantic information of the third feature data through the self - correlation reconstruction module, and capture the local and global correlations for the support - query feature pairs in each semantic level; use the local and global correlations to perform feature correlation reconstruction on the third feature data to obtain the fourth feature data with global correlations.

[0103] Step 106: Input the fourth feature data into a preset initial context - aware model for model training to obtain the target context - aware model.

[0104] Use the fourth feature data as training data and train the initial context - aware model using conventional model training techniques.

[0105] Use a large amount of annotated text data to train the initial context - aware model so that it can understand the context information of the text data and extract text segments that meet specific conditions.

[0106] Model evaluation and optimization are important steps to ensure technical performance. Evaluate the trained model using test data and optimize the model according to the evaluation results. This includes learning rate strategies and tuning, regularization techniques, etc., to avoid overfitting and improve the generalization ability of the model.

[0107] Figure 2 It is a schematic diagram of the interaction scenario of a perception model provided by an embodiment of the present invention.

[0108] As Figure 2 shown, in a typical interaction scenario, when a user interacts with a large language model, the large language model uses an enhanced large language model to enhance the interaction information using additional information and historical information. The enhanced information helps the large language model understand and parse the user's query information, comprehensively improving the effect and experience of the interaction.

[0109] The main innovation points of this solution are reflected in the following aspects:

[0110] 1) Multi-scale channel attention mechanism. In the implementation of existing technologies, feature fusion often ignores the impact of scale changes across different levels. In this model, an attention-induced cross-level fusion mechanism is introduced to effectively fuse features across different levels and use multi-scale information to mitigate the impact of scale changes. The innovation point is the introduction of a dual-branch structure. One branch uses global average pooling to obtain global context, and the other branch maintains the original feature size to obtain local context.

[0111] 2) Attention-induced cross-level fusion module technology. In the implementation of existing technologies, the adaptability of the model to targets of different scales is limited. In this model, the multi-scale channel attention mechanism fuses features at different levels by introducing an attention-induced cross-level fusion mechanism, enhancing the adaptability of the model to targets of different scales. The innovation point is to emphasize large objects with global distribution and local context, avoiding the neglect of small objects.

[0112] 3) Dual-branch global context module strategy. In the implementation of existing technologies, feature fusion often lacks the effective use of global context information. In this model, the dual-branch global context module uses rich global context information to enhance cross-level fusion features. It works through two branches. One branch captures local features through convolution operations, and the other branch captures global features through average pooling. The innovation point is the design of the two branches, effectively combining local and global information.

[0113] 4) Self-correlation reconstruction module strategy. In the implementation of existing technologies, the capture of context semantic information for feature pairs is insufficient. In this model, the self-correlation reconstruction module is used to capture the context semantic information of feature pairs and capture local and global correlations for each semantic level. The innovation point is to capture correlations for each semantic level, enhancing the richness of semantic information.

[0114] In addition, this model has strong generalization ability and transfer ability. It can be migrated to different occasions for different business scenarios, and scene adaptation and integration can be achieved at a relatively low cost. This model adopts a design method with high cohesion and low coupling, achieving high cohesion in each sub-module and being compatible with the extension of other modules or functions, resulting in excellent compatibility and strong scalability.

[0115] In summary, in the embodiments of the present invention, feature keywords and the number of feature levels are set according to the domain features of the original text sample; multi-level feature capture is performed on the original text sample according to the feature keywords and the number of feature levels to obtain first feature data; the first feature data is verified according to a preset verification rule, and the first feature data that passes the verification is used as second feature data; feature-level fusion is performed on the second feature data to obtain third feature data; feature correlation reconstruction is performed on the third feature data to obtain fourth feature data with global correlation; the fourth feature data is input into a preset initial context-aware model for model training to obtain a target context-aware model. Through the multi-scale channel attention mechanism and the attention-induced cross-level fusion module, this solution can effectively fuse features at different levels and enhance the adaptability of the model to targets at different scales. Moreover, the dual-branch global context module works through two branches. One branch captures local features through convolution operations, and the other branch captures global features through average pooling, effectively combining local and global information. The feature enhancement module is used to enhance the feature expression of feature pairs. By inputting and processing the feature pairs, the expression ability of the feature pairs is enhanced. The self-correlation reconstruction module is used to capture the context semantic information of feature pairs and capture local and global correlations for each semantic level, improving the richness of semantic information.

[0116] Figure 3 is a structural block diagram of a training device for a context-aware model provided by an embodiment of the present invention. As Figure 3 shown, the device 200 includes:

[0117] A feature setting module 201, configured to set feature keywords and the number of feature levels according to the domain features of the original text sample;

[0118] A feature capture module 202, configured to perform multi-level feature capture on the original text sample according to the feature keywords and the number of feature levels to obtain first feature data;

[0119] A verification module 203, configured to verify the first feature data according to a preset verification rule, and use the first feature data that passes the verification as second feature data;

[0120] A fusion module 204, configured to perform feature-level fusion on the second feature data to obtain third feature data;

[0121] A reconstruction module 205, configured to perform feature correlation reconstruction on the third feature data to obtain fourth feature data with global correlation;

[0122] A training module 206, configured to input the fourth feature data into a preset initial context-aware model for model training to obtain a target context-aware model.

[0123] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0124] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0125] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0126] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A method for training a context-aware model, characterized in that: The method comprises: Set feature keywords and feature levels according to the domain characteristics of the original text sample; According to the feature keywords and the feature level number, multi-level feature capture is performed on the original text sample to obtain first feature data; Verify the first feature data according to a preset verification rule, and use the first feature data that passes the verification as the second feature data; Performing feature-level fusion on the second feature data to obtain third feature data; Reconstructing the third characteristic data based on characteristic correlation to obtain fourth characteristic data with global correlation; The fourth feature data is input into a preset initial context-aware model for model training to obtain a target context-aware model.

2. The method according to claim 1, characterized in that The step of performing multi-level feature capture on the original text sample according to the feature keywords and the feature level number to obtain first feature data includes: According to the feature keywords and the feature level number, the original text sample is subjected to convolution operation to capture local features, and the global features are captured by average pooling; The local feature and the global feature are combined to obtain first feature data.

3. The method according to claim 1, characterized in that The performing feature-level fusion on the second feature data includes: Aggregating local information of the second feature data using deep convolution, capturing multi-scale context information of the second feature data using multi-branch deep convolution, and simulating the relationship between different channels in the second feature data using 1×1 point-by-point convolution to obtain first relationship data; Using the first relation data as the weight of convolutional attention; extracting low-level features and high-level features from the second feature data; The low-level features and the high-level features are cross-layer fused using the weights to obtain cross-layer fused features.

4. The method according to claim 1, characterized in that: The step of reconstructing the feature correlation of the third feature data to obtain fourth feature data having global correlation includes: Obtaining a support-query feature pair from the third feature data; Obtaining contextual semantic information of the third feature data through an autocorrelation reconstruction module, and capturing local and global correlations for the support-query feature pairs in each semantic level; The local and global correlations are used to reconstruct the feature correlation of the third feature data to obtain fourth feature data with global correlation.

5. The method according to claim 1, characterized in that The original text sample is data in the field of data security, and the setting of feature keywords and feature levels according to the field characteristics of the original text sample includes: The characteristic levels of data security risks are set to three levels, among which the first-level characteristics include: data carrier risk, transmission risk, storage risk, application risk, sharing risk, destruction risk, etc. A plurality of secondary features are set for each of the primary features; the secondary features of the storage risk include: physical storage risk and software storage risk; A plurality of third-level features are set for each of the second-level features; the third-level features of the software storage risk include software vulnerability risk, software version risk, software access risk, and software original storage risk.

6. A training device for a context-aware model, characterized in that: The device comprises: A feature setting module is used to set feature keywords and feature levels according to the domain characteristics of the original text sample; A feature capture module, used for performing multi-level feature capture on the original text sample according to the feature keywords and the feature level number to obtain first feature data; A verification module, used to verify the first feature data according to a preset verification rule, and use the first feature data that passes the verification as the second feature data; A fusion module, used for performing feature-level fusion on the second feature data to obtain third feature data; A reconstruction module, used for reconstructing the third characteristic data based on characteristic correlation to obtain fourth characteristic data with global correlation; The training module is used to input the fourth feature data into a preset initial context-aware model for model training to obtain a target context-aware model.

7. The device according to claim 6, characterized in that The feature capture module is specifically used for: According to the feature keywords and the feature level number, the original text sample is subjected to convolution operation to capture local features, and the global features are captured by average pooling; The local feature and the global feature are combined to obtain first feature data.

8. The device according to claim 6, characterized in that The fusion module is specifically used for: Aggregating local information of the second feature data using deep convolution, capturing multi-scale context information of the second feature data using multi-branch deep convolution, and simulating the relationship between different channels in the second feature data using 1×1 point-by-point convolution to obtain first relationship data; Using the first relation data as the weight of convolutional attention; extracting low-level features and high-level features from the second feature data; The low-level features and the high-level features are cross-layer fused using the weights to obtain cross-layer fused features.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the context-aware model training method as described in any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method of a context-aware model as described in any one of claims 1 to 5.