Method and device for studying and judging data security

By performing semantic analysis and large language model processing on security assessment texts, the limitations of existing technologies, such as single data security assessment rules and manual judgment, are solved, achieving efficient and intelligent security assessment and improving the accuracy and adaptability of data security management.

CN121350853APending Publication Date: 2026-01-16CHINA MOBILE GROUP SHANDONG +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302642.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

In existing technologies, data security assessment rules are singular and rely on manual setting, resulting in limitations and inadequacies. They cannot effectively address diverse security needs and rapidly changing threats, and manual judgment is subject to subjectivity and inefficiency.

Method used

By acquiring the text of security assessment questions, performing text semantic parsing, determining the security assessment requirements and data source information, matching target data from preset multi-source security data, applying a large language model for processing, and generating security assessment answers, a highly automated and intelligent security assessment is achieved.

Benefits of technology

It improves the real-time performance and accuracy of security assessments, reduces the limitations of human judgment, enables rapid processing of massive amounts of information, dynamic response to security threats, and generation of reliable security assessment conclusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121350853A_ABST
    Figure CN121350853A_ABST
Patent Text Reader

Abstract

The invention discloses a data security research and judgment method and device, which are used for improving the effectiveness of data security research and judgment. According to the scheme, the method comprises the steps of obtaining a security study and judgment question text; executing text semantic analysis on the security research and judgment problem text to obtain security research and judgment demand information and data source information of security research and judgment problem text semantic representation; determining target security data matched with the data source information from preset multi-source security data; performing processing matched with the security research and judgment demand information on the target security data to obtain processed security data; and based on the processed security data, generating a security research and judgment answer in response to the security research and judgment question text.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data security, and in particular to a data security research and judgment method and device. BACKGROUND

[0002] In the field of data security, with the continuous development of informatization, networking and intelligentization, data security is challenged in many aspects. The diversification of data and the complex security research and judgment requirements make it difficult for a single data security research and judgment rule to meet the existing needs.

[0003] If the security research and judgment rule is set by a technical personnel, the set rule is affected by the subjective experience of the technical personnel and has limitations. In addition, the application scenarios of data security research and judgment are wide, and the security research and judgment rule effective for a specified scenario cannot be effectively applied to another scenario.

[0004] If the effectiveness of data security research and judgment is improved, it is the technical problem to be solved by the present application. SUMMARY

[0005] The purpose of the embodiments of the present application is to provide a data security research and judgment method and device to improve the effectiveness of data security research and judgment.

[0006] In a first aspect, a data security research and judgment method is provided, comprising: obtaining a security research and judgment question text; performing text semantic analysis on the security research and judgment question text to obtain security research and judgment demand information and data source information semantically represented by the security research and judgment question text; determining target security data matched with the data source information from preset multi-source security data; performing processing matched with the security research and judgment demand information on the target security data to obtain processed security data; generating a security research and judgment answer responding to the security research and judgment question text based on the processed security data.

[0007] In a second aspect, a data security research and judgment device is provided, comprising: an obtaining module that obtains a security research and judgment question text; an executing module that performs text semantic analysis on the security research and judgment question text to obtain security research and judgment demand information and data source information semantically represented by the security research and judgment question text; a determining module that determines target security data matched with the data source information from preset multi-source security data; a processing module that performs processing matched with the security research and judgment demand information on the target security data to obtain processed security data; generating a security judgment answer responding to the security judgment question text based on the processed security data.

[0008] In a third aspect, an electronic device is provided, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor, which, when executed by the processor, implements the steps of the method of the first aspect.

[0009] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program, which, when executed by a processor, implements the steps of the method of the first aspect.

[0010] In a fifth aspect, a computer program product is provided, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of the method of the first aspect.

[0011] In the embodiments of the present application, first, a security judgment question text is acquired; then, text semantic analysis is performed on the security judgment question text to obtain security judgment requirement information and data source information of semantic representation of the security judgment question text; next, target security data matching the data source information is determined from preset multi-source security data; subsequently, processing matching the security judgment requirement information is performed on the target security data to obtain processed security data; finally, a security judgment answer responding to the security judgment question text is generated based on the processed security data. Through the scheme provided in the embodiments of the present application, effective analysis can be performed on the security judgment question text in the security judgment scenario, so as to determine the security judgment requirement information and the data source information. In this way, in the subsequent steps, the target security data can be effectively acquired and processing matching the security judgment requirement information can be performed. Thus, the security data obtained through the processing can meet the semantic requirement of the security judgment question text, and the security judgment answer can effectively respond to the security judgment question, improving the effectiveness of data security judgment. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application and illustrate embodiments of the present application and its description, which serve to explain the present application, but do not constitute improper limitations on the present application. In the drawings: Figure 1a is one of flow diagrams of a method of data security judgment according to an embodiment of the present application; Figure 1b is another of flow diagrams of a method of data security judgment according to an embodiment of the present application; Figure 2ais a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 2b is a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 3 is a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 4a is a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 4b is a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 5a is a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 5b is a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 6a is a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 6b is a flowchart of a method for data security research and judgment according to an embodiment of the present application; Figure 7 is a structural diagram of a device for data security research and judgment according to an embodiment of the present application. DETAILED DESCRIPTION

[0013] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application. The drawing numbers in the present application are only used to distinguish each step in the scheme, and are not used to limit the execution order of each step, which is subject to the description in the specification.

[0014] In the field of security research and judgment, if relying on manual collection and analysis of information, it is difficult to quickly and accurately extract effective information when facing massive data, which easily leads to the lag of research and judgment results. Specifically, there are defects such as slow information processing speed, low efficiency, limited processing capacity, etc.

[0015] In addition, this artificial processing method relies on the experience and intuition of experts for judgment, and this subjective method can be affected by individual cognitive limitations, thereby affecting the accuracy and comprehensiveness of the safety judgment. Specifically, the judgment result can be affected by personal bias, emotion, and cognitive limitations. Moreover, human judgment is usually limited by available information and cannot obtain or process all relevant data, affecting the comprehensiveness and accuracy of the judgment. In addition, the knowledge of experts and decision-makers may not be updated in a timely manner, which adapts to the changing technology, market and social environment. As can be seen, due to the limitations of human judgment, the judgment result often cannot fully predict potential risks and uncertainties, leading to increased decision-making risk. Moreover, different experts or decision-makers may have different judgment standards and preferences, which can lead to poor consistency and repeatability of the judgment result.

[0016] If fixed safety judgment rules are used to replace artificial processing methods, although the consistency of the judgment result and the processing efficiency can be improved, such fixed safety judgment rules often cannot be updated in a timely manner and cannot respond quickly to new security threats and challenges, making it difficult to meet the rapidly changing security needs. Once new risks appear in the scene, the fixed safety judgment rules are difficult to respond effectively, which may lead to delayed management and control, and the security vulnerability is expanded. In addition, the fixed safety judgment rules have various application limitations and cannot be flexibly applied to different data streams.

[0017] In order to solve the problems existing in the related art, the embodiments of the present application provide a data security judgment method, as shown in Figure 1a The method comprises the following steps: S11: Obtain a safety judgment question text.

[0018] The above safety judgment question text can be a text entered by a user, or it can also be a text generated based on human-computer interaction.

[0019] In actual application, the safety judgment question text can be obtained in various ways. For example, based on the dialogue interface of the large language model, the safety judgment question text is obtained in the human-computer interaction scene. The safety judgment question text can be related text expressing the user's safety judgment demand, which can be continuous dialogue or multiple separated texts related to safety judgment extracted from continuous dialogue.

[0020] S12: Perform text semantic analysis on the safety judgment question text to obtain safety judgment demand information and data source information of semantic representation of the safety judgment question text.

[0021] The present step can be implemented through natural language processing technology. Specifically, a large language model can be used to perform text semantic analysis on the security research problem text, so as to extract the security research demand information and data source information expressed by the security research problem text from the semantic dimension.

[0022] The security research demand information refers to the demand conditions related to security research, which can be used for security research on data. The data source information refers to the data source that needs to perform security research. The data source information can include the form of the data source, such as stream data or single file data. The data source information can also include the type, key field, structure, and other attribute information of the data.

[0023] S13: Determine target security data matching the data source information from the preset multi-source security data.

[0024] In this step, the target security data that meets the above data source information is determined from the preset multi-source security data. The preset multi-source security data can be obtained and preprocessed in advance.

[0025] For example, referring to Figure 1b , first, the acquisition module collects multi-source data from multiple different business systems. This can include structured (such as database table data), semi-structured (such as JSON or XML files), and unstructured data (such as log files or text content). Historical security business data is collected to establish a data management system.

[0026] The data sources can include, for example, a situational awareness platform, an asset management platform, a vulnerability management platform, a log audit management platform, and other security business systems.

[0027] Specifically, these data can be extracted through API (Application Programming Interface), data synchronization tools, or ETL (Extract, Transform, Load) processes, ensuring data integrity and timeliness.

[0028] Subsequently, to improve the quality of the collected raw data, cleaning processing can be performed. The cleaning processing can include deleting noise data, filling missing values, and correcting abnormal values. Standardization processing can also be performed to unify the data into a standardized format suitable for large language model analysis. Through this step, the acquired data is corrected to obtain a cleaned data set.

[0029] Optionally, key feature selection and generation can also be performed on the data, thereby extracting variables of high value for security judgment, retaining key features and reducing the amount of data, laying a high-quality data foundation for subsequent large language model analysis.

[0030] For example, based on the analysis module, key entities (such as device names, vulnerability identifiers), relationships (such as causal relationships, dependency relationships), and other key information related to security judgment are extracted from the collected data through natural language processing technology. These information can be represented as a knowledge graph or an embedded vector after natural language processing, which can facilitate the large language model to better understand the semantic content of the data. In addition, to improve the correlation analysis capability of cross-system data, concept mapping technology can also be used to uniformly identify and represent the same or similar business entities from different systems. In this way, the pre-collected multi-source data can be pre-processed into feature vectors with rich semantics and operability, providing support for data matching.

[0031] In this step, the language understanding and reasoning capabilities of the large language model can be used to determine the target security data that matches the data source information from the pre-set multi-source security data. Specifically, multi-source security data can be used for deep correlation analysis to determine target security data that matches the data source information.

[0032] S14: Perform processing on the target security data that matches the security judgment requirement information to obtain processed security data.

[0033] In this step, based on the security judgment requirements, the large model can be used to quickly locate security judgment problems and identify risk factors. For example, when facing potential security threats, the time series and behavior patterns of the target security data can be analyzed through the large model to mine hidden risk characteristics. Based on the analyzed security risks, further analysis of targeted solutions can be performed, such as generating repair guidance or optimization suggestions. This solution utilizes the dynamic adaptability and real-time decision-making capabilities of the model to ensure efficient response in complex security scenarios, generating reliable security judgment conclusions. Further, based on the security judgment conclusions and security judgment requirement information, efficient processing of the target security data can be performed to obtain processed security data.

[0034] Optionally, the processing of the target security data can also include sending an alarm to the relevant business system, updating the security policy, or triggering remedial measures.

[0035] S15: Based on the processed security data, generate a security judgment answer that responds to the security judgment problem text.

[0036] In this step, a security research and judgment answer is generated based on the processed security data, which is used to respond to the security research and judgment question text. In actual application, it can be sent to the user who raises the security research and judgment question in the form of a dialogue. In addition, the evaluation and feedback information of the user on the analysis result can also be recorded and used as the basis for optimizing the model to continuously improve the accuracy of model analysis and the timeliness of feedback. At the same time, this closed-loop mechanism can ensure that the system realizes dynamic adjustment and performance improvement in continuous operation, so that the model can adapt to the changing security environment and provide more reliable security research and judgment support for users.

[0037] The scheme provided by the embodiments of the present application is applied to the field of security research and judgment, can realize high automation and high intelligence, and effectively improve the real-time performance and accuracy of security research and judgment. Through intelligent security research and judgment, massive information can be processed faster, and the limitations of human judgment can be reduced to realize dynamic, real-time and systematic security research and judgment. Through the scheme provided by the embodiments of the present application, the security research and judgment question text can be effectively analyzed in the security research and judgment scene, so as to determine the security research and judgment demand information and the data source information. In this way, the target security data can be effectively obtained and the processing matched with the security research and judgment demand information can be performed in the subsequent steps. Therefore, the security data obtained by processing can meet the semantic requirements of the security research and judgment question text, and the security research and judgment answer can effectively respond to the security research and judgment question, thereby improving the effectiveness of data security research and judgment.

[0038] Next, the present scheme will be further described in combination with an example.

[0039] In the security research and judgment application scene, first, structured, semi-structured and unstructured data are collected and integrated from different business systems. The collected raw data is cleaned and standardized to be converted into a unified format, from which key data is extracted. Among them, natural language processing technology can be used to extract entities, relationships and other meaningful information in the business system, and convert them into knowledge graph or embedded vector representation, which is beneficial to the full understanding and analysis of related business entities among different system data by establishing cross-system concept mapping. Then, the language understanding and reasoning ability of the large model is applied to determine the target security data matched with the data source information from the preset multi-source security data, and perform processing matched with the security research and judgment demand information on the target security data to obtain processed security data. Among them, intelligent scheduling can be used to quickly locate the problem and provide intelligent detection functions such as cause analysis and solution, so as to optimize the processing effect of data. Then, according to the analysis conclusion, corresponding feedback processing is performed, and the security research and judgment answer effectively responds to the security research and judgment question text, realizing efficient and intelligent security research and judgment.

[0040] In the scheme provided by the embodiments of the present application, the text semantic analysis is applied to improve the effectiveness of security research and judgment. The matching degree between different data and text is accurately calculated through the analysis of security research and judgment problem text, the extraction of key concepts, and the construction of semantic network, and the like, the multi-source data correlation analysis is realized, the comprehensiveness of the target security data obtained by identification is improved, and the accuracy and efficiency of analysis are improved.

[0041] In addition, the scheme provided by the embodiments of the present application performs semantic analysis on security research and judgment problem text, and improves the effectiveness of security research and judgment response from two aspects of security research and judgment demand information and data source information. It can be effectively applied to multi-source security data with complex structure, has scalability and stability, and ensures the efficiency and availability of the scheme in complex application scenarios.

[0042] Compared with the security research and judgment mode realized by manual processing, the scheme significantly improves the processing efficiency of research and judgment. Errors may occur during manual processing, while intelligent processing can reduce these errors and ensure accuracy. The self-learning ability of the large model can continuously learn and optimize according to feedback and new data, improve the quality of research and judgment, and improve the accuracy of data analysis. In practical application, the scheme not only helps to improve the level of data security management, but also provides strong guarantee for data development.

[0043] Based on the scheme provided in the above embodiments, as shown in Figure 2a Before step S13, before determining the target security data matched with the data source information from the preset multi-source security data, the following steps are further included. S21: Obtain security data to be processed from a plurality of data sources.

[0044] In this step, a plurality of security data to be processed in various forms can be obtained from a plurality of data sources. For example, it can be structured, semi-structured and unstructured data. Optionally, the security data to be processed is preprocessed to improve data quality.

[0045] S22: Perform correlation analysis on the security data to be processed from different data sources to obtain correlation information between the security data.

[0046] In this step, the security data to be processed from different data sources is subjected to correlation analysis. For example, the security data to be processed can be subjected to feature extraction and vectorization processing, and the correlation information between different security data is determined by comparing the similarity of vector features. Alternatively, the data containing the same keywords can be determined as the data having correlation by means of data keyword retrieval. Alternatively, the data with similar meanings can be determined as the data having correlation by means of semantic analysis.

[0047] S23: constructing the multi-source security data based on the association information and the security data to be processed from different data sources.

[0048] Based on the association information, the data with the association is integrated to obtain the multi-source security data with the association structure. Through the scheme provided in the embodiments of the present application, the security data obtained in advance is processed, the association between the security data obtained from different sources can be effectively highlighted, which is beneficial to data matching of the model and improves the comprehensiveness of the target security data obtained in the subsequent step.

[0049] In the scheme provided in the embodiments of the present application, data from multiple security systems can be collected. In actual application, each system or module can be good at processing a certain type of problem, and through integration, the advantages and data assembly of each system can be fully utilized to improve the overall question answering capability. In the process of data processing, useful information can be extracted from a large amount of data, which is beneficial to more fully answering the user's question.

[0050] The scheme provided in the embodiments of the present application can be implemented through natural language processing technology, which can specifically include steps of entity recognition, relationship extraction, event extraction, etc. By analyzing the relationship between different data, the large model can better understand the data and improve the security research and judgment accuracy.

[0051] For example, for the data accuracy problem, the large model can more accurately understand the semantics of the text through text semantic analysis. In a complex intelligent question answering system, by obtaining the security research and judgment problem text and performing analysis, various security requirements can be extracted. Then, by matching the target security data and performing processing, comprehensive data security research and judgment can be effectively realized, which is beneficial to improving the performance and accuracy of the question answering system and better serving the user.

[0052] Next, the present scheme will be further described in combination with an example.

[0053] Optionally, the scheme provided in the embodiments of the present application can be executed by a data security research and judgment system, which can include a data integration and preprocessing module, a knowledge extraction and representation module, an association analysis and reasoning module, and an intelligent question answering module. The data integration and preprocessing module is used to collect multiple security systems, the knowledge extraction and representation module is used to extract useful information from a large amount of data, the association analysis and reasoning module is used to analyze the relationship between different information, and the intelligent question answering module is used to receive various security requirements in a complex intelligent question answering system, process different types of problems and generate answer responses. The structure of the system is shown in Figure 2b .

[0054] The data integration and preprocessing module extracts key data from multiple security business platforms such as the situational awareness platform, asset security management platform, vulnerability management platform, log audit management platform, data security management platform, identity and access security management platform, and security operation platform in real-time or at regular intervals through API interface, data synchronization tool, or ETL (Extract, Transform, Load) process. These data include but are not limited to: Situational awareness data: attack events, threat intelligence, network traffic, etc. Asset data: device list, software version, patch status, etc. Vulnerability data: vulnerability details, impact range, repair status, etc. Log audit data: user behavior, system events, abnormal alarms, etc. Data security data: sensitive data distribution, data flow record, data leakage event, etc. Identity and access data: user account, permission assignment, login behavior, etc. Security operation data: emergency response record, security policy, security event statistics, etc.

[0055] The data integration and preprocessing module can be used to collect data from different sources, which may be structured or unstructured. The purpose of data integration is to create a unified data portal to support data analysis. This includes extracting and integrating data from various data sources; then converting or mapping data into a unified format to ensure compatibility of data from different sources; loading the converted data into the target system; finally identifying and correcting errors and inconsistencies in the data to ensure data accuracy and reliability.

[0056] The preprocessing module is used to process the raw data obtained to improve data quality and prepare it for subsequent data analysis and model training. This module can be used to remove noise data, handle missing values, correct outliers, standardize or normalize data, etc., so that the processed data is suitable for specific algorithms or analysis needs. In the processing process, identify and select variables that best represent data characteristics to simplify the model and reduce the risk of overfitting. In addition, new features can be created or existing features can be transformed to provide more information to the model. Reduce the dimensionality of the data through techniques such as principal component analysis (PCA). Ensure the consistency, accuracy and availability of data from the original source to the final analysis stage. Through these processes, the reliability of data analysis and machine learning model results can be significantly improved.

[0057] For the acquisition and preprocessing of security data, the following methods can be used: First, determine which data needs to be collected and the sources of each type of data. Collect data from data sources such as security systems, databases, and other data sources. Clean the collected data to ensure its accuracy and quality. The data cleaning process includes detecting and correcting outliers, missing values, and duplicates. Then, integrate the cleaned data into a complete dataset. Finally, support the establishment of a data management system to manage data, including storage, backup, security, and access control, to better understand and analyze data.

[0058] Optionally, the data collection steps are as follows: Requirement analysis: Clearly define the purpose and requirements of data collection, determine the types, sources, and scope of data to be collected.

[0059] Data source selection: Select appropriate data sources based on requirements, such as databases, files, online interfaces, logs, etc.

[0060] Data access: Connect data sources to the data collection system, involving API calls, data import / export, web scraping, etc.

[0061] Data extraction: After establishing data connections, extract the required data fields according to the predetermined logic.

[0062] Data conversion: Convert the extracted data to meet the needs of subsequent processing, such as timestamp processing, data type conversion, etc.

[0063] Data loading: Load the converted data into data storage warehouses, such as data warehouses, data lakes, etc.

[0064] Through the above data collection steps, data can be integrated across various security systems. Actual security data often contains a large amount of noise and outliers, which may come from data collection errors, or outliers may be potential security threats. In order to generate and process data-based research and judgment, it is necessary to process raw data into high-quality data to lay a solid foundation for subsequent security data analysis and application.

[0065] Optionally, the data cleaning steps are as follows: Data quality check: Check the collected data for completeness, consistency, accuracy, etc.

[0066] Missing value processing: Handle missing values in data, which can be filled, deleted, or processed using interpolation methods.

[0067] Outlier processing: Correct outliers, which may be caused by data entry errors or unreasonable data.

[0068] Duplicate data processing: Detect and process duplicate data, which can be deleted or merged.

[0069] Data standardization: Standardize the data, such as uniform unit, normalization, etc., to eliminate the influence of dimension between data.

[0070] Feature engineering: According to the needs of data analysis, generate new features or transform existing features to improve the performance of the model.

[0071] Data verification: Verify the cleaned data to ensure the effectiveness of data cleaning.

[0072] Based on the above embodiments, the scheme can be selected as shown in Figure 3 Before step S22, that is, before performing correlation analysis on the security data from different data sources to obtain the correlation information between the security data, it further includes: S31: Perform anomaly detection on the security data from multiple data sources by the Isolation Forest algorithm; S32: Mark the security data detected as abnormal as an abnormal state; In step S22, the correlation analysis on the security data from different data sources is performed to obtain the correlation information between the security data, including: S33: Perform correlation analysis on the security data from different data sources based on the security data marked with an abnormal state to obtain the correlation information between the security data.

[0073] In actual operation, the process of data collection and cleaning needs to be iterated and optimized continuously to adapt to the changes of data characteristics and analysis requirements. Among them, the security data anomaly detection can be realized by the Isolation Forest algorithm (Isolation Forest). For example, it is used to process part of the network traffic data of our third-party platform flow product, each record contains timestamp, source IP (Internet Protocol) address, target IP address, protocol type, entry port, exit port and traffic size, etc. The detection target is to identify the abnormal traffic in these records, which can help to discover possible attack behavior in time.

[0074] The specific steps based on the Isolation Forest algorithm are as follows: First, a feature and a split point are randomly selected. A feature is randomly selected from the security data on which anomaly detection needs to be performed, and a split point is randomly selected. This random selection process is repeated to build multiple trees. Subsequently, for each data point in the data set, the path length thereof is calculated through the built trees. Then, for each data point, the average of the path lengths thereof in the multiple trees is calculated to obtain an anomaly score of the data point. According to the anomaly scores of all data points, a threshold is determined. If the anomaly score of a data point is greater than the threshold, the data point is marked as an anomaly.

[0075] Based on the Isolation Forest algorithm, normal data points can have relatively short path lengths in multiple trees through random partitioning, while anomaly data points can have long path lengths. By comparing the path lengths, the algorithm can identify anomaly points. In practice, the Isolation Forest algorithm usually needs a large number of trees (e.g., more than 1000 trees) to improve the accuracy of detection. The time complexity of the algorithm depends on the number of trees, the size of the data set, and the number of features.

[0076] Through the scheme provided by the embodiments of the present application, records with path lengths significantly higher than normal data can be identified as abnormal. Subsequently, the data marked as abnormal can be used for data correlation analysis, which is beneficial to the accurate understanding of the correlation between data by a large model, thereby improving the accuracy of data matching and the efficiency of data security research and judgment.

[0077] Based on the scheme provided by the above embodiments, as shown in Figure 4a In the step S13, the target security data matched with the data source information is determined from the preset multi-source security data, which includes the following steps. S41: Perform vector similarity calculation on the preset multi-source security data and the security research and judgment requirement information to obtain first security data matched in vectors.

[0078] In this step, the feature extraction can be performed on the preset multi-source security data and the security research and judgment requirement information respectively to obtain each security data in vector form and the security research and judgment requirement information in vector form. Then, the vector similarity between the feature vector of the security research and judgment requirement information and the feature vector of each security data is calculated respectively. The first security data matched in vectors is determined as the security data whose vector similarity is greater than a preset value or the preset number of security data with the maximum vector similarity. Optionally, the feature extraction is performed by a large model to convert the above security data and security research and judgment requirement information into vector form.

[0079] S42: Perform semantic matching judgment on the preset multi-source security data and the security research and judgment requirement information to obtain second security data matched in semantics.

[0080] In this step, semantic analysis can be performed on the security judgment demand information first, and second security data that is semantically matched is queried from the preset multi-source security data based on the semantic analysis result. The semantic analysis can specifically include semantic analysis of keywords, sentiment intention, key entities, etc. Through semantic matching, the second security data that is semantically matched with the security judgment demand information is filtered from the preset multi-source security data.

[0081] Referring to Figure 4b In this step, the target demand can be further determined based on the pre-acquired security judgment demand information, and then the feature vectors in the text are extracted for the security judgment demand information, the matching set is matched, and the set containing the security data is obtained. The security data in the set is semantically matched with the security judgment demand information. Subsequently, according to the matching set, a context semantic matching relationship result is generated, and the second security data is determined.

[0082] S43: Determine the target security data based on the first security data and the second security data.

[0083] According to the first security data and the second security data filtered by the vector similarity and the semantic matching in the above steps, the target security data is determined in this step. Optionally, the intersection of the first security data and the second security data is determined as the target security data. Alternatively, based on a preset statistical rule, the first security data and the second security data are respectively filtered to obtain the target security data containing at least part of the first security data and at least part of the second security data. Or, the first security data and the second security data are collectively used as the target security data.

[0084] Next, the present scheme will be further described by examples. The present scheme can be executed by the knowledge extraction and representation module in the above Figure 2b .

[0085] In the extraction method of the security data, the security data filtering can be first performed according to the pre-defined semantic matching and the correspondence relationship between the actual business fields. Then, the position information of the semantic features of each security data between systems is determined. Subsequently, the index data on each position information is extracted. The security data to be extracted (such as the statistical period of the financial judgment) and the index data on each position information, the index code of each target marker character are stored to obtain an index data storage table, thereby realizing the extraction of the security data.

[0086] In the processing method of the security judgment demand information, the user can convey the security judgment demand information to the electronic device through human-computer interaction, so that the machine obtains the target demand. The user inputs the security judgment problem text containing the security judgment demand information and the data source information (such as the demand rule of the audit file processing), which can be used for cross-system calling of the data source.

[0087] Optionally, the security research and judgment problem text is identified to obtain security research and judgment requirement information and collected and stored. Subsequently, an SQL (Structured Query Language) statement is generated according to the security research and judgment requirement information to call a data source. In the extraction steps of the security research and judgment requirement information and the data source information, various technologies and algorithms can be used to enable the machine to understand natural language and identify keywords in order to better meet the needs of users. For example, the following methods can be used to extract information from the security research and judgment problem text: Natural language processing (NLP): NLP technology is used to understand and parse text, which can include word segmentation, part-of-speech tagging, named entity recognition, sentiment analysis, etc.

[0088] Intention recognition: based on understanding the text, the user's intention is identified, which can use machine learning models such as classifiers or sequence labeling models.

[0089] Context understanding: in order to more accurately obtain the requirements, the security research and judgment problem text is analyzed in combination with relevant context information, such as time, location, user's historical behavior, etc.

[0090] Requirement abstraction and representation: the collected information is abstracted into a feature vector or other form that is easy for the model to process, such as a feature vector, a label, or a classification.

[0091] Optionally, after analyzing the security research and judgment problem text, the model can be fed back and continuously optimized based on the analysis results to improve the performance of the model.

[0092] Optionally, feature extraction is performed by a machine learning model to convert the text into a quantifiable feature vector. This can be achieved through techniques such as bag-of-words, TF-IDF (Term Frequency-Inverse Document Frequency), word embedding, etc.

[0093] In practical applications, the content of the pre-set multi-source security data is often complex, with complex structures and important and non-important information content. In this scheme, feature extraction can be performed on the pre-set multi-source security data to perform vector similarity calculation or semantic matching judgment. Optionally, security data cannot be effectively extracted by regular expression parsing or key feature method, which may result in low utilization of security data value, inaccurate data analysis, and low data analysis level. For complex security data special field content, the bag-of-words text feature representation method can be used to extract features from complex content, thereby accurately extracting high-value data.

[0094] The parsing method based on the bag-of-words text feature representation method can first remove punctuation, numbers, and special characters from the security data content, and then perform word segmentation, stem extraction, and morphological reduction operations to unify different forms of words into their basic forms, which helps to eliminate noise and irrelevant information for security data content with high complexity, making the data more clear and standardized. Subsequently, a vocabulary table is constructed. Specifically, non-repeated words are extracted from the security data to construct a vocabulary table. The size of the vocabulary table can be controlled by setting a word frequency threshold, thereby filtering out words with too low frequency of occurrence, so as to realize centralized display of high-value words in the security data content, making the subsequent data reading device of the security operation workflow more effectively understand the key words and theme information in the text.

[0095] The feature extraction using the bag-of-words (BOW) model can be implemented based on the following features: Term List: Break down the text into words or phrases, creating a term list that does not consider the order of words.

[0096] Term Frequency (TF): Calculate the number of times each term appears in the text.

[0097] Inverse Document Frequency (IDF): Measures the importance of a term across all documents, often used to reduce the importance of common words.

[0098] TF-IDF Vector: Based on the above TF and IDF, a TF-IDF vector is generated by combining it, creating a TF-IDF weight for each term.

[0099] Where TF-IDF is similar to the bag-of-words model, TF-IDF not only considers the frequency of words, but also considers the importance of words.

[0100] Optionally, semantic matching judgment is achieved by understanding and matching semantic information in the context of the text. Specifically, it can involve understanding the meaning of words, phrases, sentences, or documents, as well as their relationships. The purpose of matching is to accurately understand and interpret the user's intent and expression, especially in the context of file processing and rule generation.

[0101] Semantic matching judgment specifically can include context semantic matching judgment, and its key elements can include: Word Sense Disambiguation: In different contexts, the same word can have different meanings. Contextual semantic matching needs to be able to identify and distinguish these different meanings.

[0102] Phrase comprehension and matching: A phrase may consist of multiple words, which together have a specific meaning. For example, the word "apple" in "buy apples" and "Apple phone" have different meanings.

[0103] Sentence intent recognition: Understanding the overall intent and goal of a sentence, which may involve sentiment analysis, subject-verb-object (SVO) structure analysis, etc.

[0104] Document-level contextual understanding: When processing long texts, it is necessary to understand the context of the entire document in order to correctly match and understand the content of specific parts.

[0105] Utilizing contextual information: Use contextual information to improve matching accuracy, such as time, location, and people.

[0106] The solution provided by the embodiments of this application can perform feature extraction on preset multi-source security data and security assessment requirement information, and perform similarity matching judgment from feature vectors and semantic expressions, thereby comprehensively determining the target security data that matches the security assessment requirement information in the preset multi-source security data, and improving the accuracy and comprehensiveness of the target security data.

[0107] Based on the solution provided in the above embodiments, optionally, the security assessment issue text includes multiple text segments based on time sequence; Among them, such as Figure 5a As shown, in step S12 above, text semantic parsing is performed on the security assessment question text to obtain the security assessment requirement information and data source information represented by the semantic representation of the security assessment question text, including: S51: Perform keyword parsing on the security assessment question text to obtain keywords corresponding to multiple text segments based on time sequence.

[0108] Optional, via Figure 2b The security analysis and reasoning module shown here performs text semantic parsing on security assessment question texts.

[0109] In this step, based on a preset matching algorithm, information mining is performed on the time-series statements during human-computer interaction to obtain keywords. These keywords often express the user's key needs or key entities.

[0110] S52: Determine the keyword similarity between different text segments based on the keywords corresponding to each segment.

[0111] In this step, the similarity between text segments is determined based on the keywords of each text scaffold. This similarity is used to represent the degree of similarity of key features between text segments, which can be used in subsequent steps to extract key information from the text and facilitate understanding user intent.

[0112] Optionally, the keyword similarity described above can also be adjusted based on the segment weight. For example, first analyze the possibility of containing the problem theme in each segment of text. Then, add the problem theme possibilities of different two segments to obtain an adjustment weight, multiply the adjustment weight with the keyword similarity of the segment based on the keyword described above to obtain the comprehensive keyword similarity corresponding to each text segment. That is, the greater the problem theme possibility between the two text segments, the greater the corresponding adjustment weight, which means that the historical problem text in the two text segments represents the theme of the same historical answer text, and the segment similarity is greater. Based on the segment weight adjustment, the reliability of the keyword similarity between the texts can be effectively improved.

[0113] S53: Perform semantic analysis on the multiple segments of text respectively to determine the security research demand information probability and data source information probability corresponding to the multiple segments of text respectively.

[0114] Specifically, the semantic feature vector of the problem text can be obtained by analyzing the question and answer data text of human-computer interaction. According to the context relationship of the text time sequence, the keywords related to the research demand or data source are obtained. The semantic relationship between the texts is analyzed by using an artificial intelligence large model to obtain the security research demand information probability and data source information probability, thereby determining what data needs to be processed according to the security research demand in the question and answer.

[0115] S54: Determine the security research demand information and data source information of the security research problem text based on the keyword similarity between the segments of text, the security research demand theme probability and the data source information probability corresponding to the multiple segments of text respectively.

[0116] In this step, the security research demand information and the data source information expressed by the user are comprehensively determined based on the security research demand information probability and the data source information probability corresponding to the multiple segments of text respectively. Optionally, the security research demand information with the highest probability is determined as the security research demand information of the security research problem text, and the data source information with the highest probability is determined as the data source information of the security research problem text.

[0117] In actual application, optionally, the user can correct the rules generated by the machine. The user can provide the file to the machine for analysis by uploading the file to be processed, which can improve the efficiency of the user's daily file processing work.

[0118] In the embodiments of the present application, the security research demand information can also be generated by the multiple segments of text in the security research problem text together. For example, the context relationship between the text segments is determined based on the keywords of different text segments to determine the data processing rule to generate the security research demand information.

[0119] wherein, for two different text reference segments, the similarity of the two segments can be determined by the following formula:

[0120] wherein, G (j,k) is the similarity between reference segment j and reference segment k, n j is the number of question texts in reference segment j, n k is the number of question texts in reference segment k, E(b x ,b y ) is the similarity of the semantic feature vector between the xth question text b x in reference segment j and the yth question text b y in reference segment k.

[0121] Based on the similarity between the above text segments, an association graph can be constructed. Further, the association graph is input into a machine learning model to identify the determinants, generate SQL statements, and match the target security data from the pre-set multi-source security data of the database according to the output analysis results.

[0122] Optionally, referring to Figure 5b , in the present scheme, the sentence attributes are mined based on the security research and judgment question texts to obtain keywords. Then, the association graph in the above step is collected, and the system resources are matched to output the reasoning result. Wherein, the target security data and the processing rule matched with the security research and judgment requirement information can be determined based on the keywords and the association graph, so as to execute the processing on the target security data according to the processing rule to output the result.

[0123] Through the scheme provided by the embodiments of the present application, the security research and judgment question texts can be deeply analyzed, and the security research and judgment requirement information and the data source information can be effectively determined, so as to effectively select the target security data according to the user's intention and execute the security research and judgment processing required by the user.

[0124] Based on the scheme provided by the above embodiments, optionally, the target security data includes user local privacy data; wherein, as Figure 6a shown, in the above step S14, the processing matched with the security research and judgment requirement information is executed on the target security data to obtain the processed security data, including: S61: executing the processing matched with the security research and judgment requirement information on the user local privacy data through the large language model deployed in the user local to obtain the processed security data.

[0125] The scheme provided by the embodiments of the present application can be applied to the private domain interaction scene, and the security of the user local privacy data can be effectively improved.

[0126] The scheme directly deploys a security industry large model in a user production environment, accesses user private data resources for intelligent question and answer. Private data refers to the historical business data accumulated among various security systems. The customized local large model is obtained by integrating user private data resources with the security large model, matching the content by combining semantic features and content similarity algorithms, realizing the precise demand of the user through one question and one answer, and generating private and precise processing rules. On the one hand, the data of the enterprise is integrated, and the intelligent question and answer technology is closely combined with the actual security demand, and the private domain interaction is also a prerequisite for our intelligent question and answer system to be different from other question and answer systems. At the same time, the security industry large model is deployed in the user's local production environment, which can greatly reduce the risk of data privacy leakage.

[0127] Optionally, to truly solve the problem of injecting file processing requirements through human-computer one question and one answer, by asking the machine in the form of a question, the machine masters the file processing requirements in the multi-round response process, supports the user to upload the original file, and finally the machine spits out a file processed based on the artificial intelligence question and answer learning rules.

[0128] Optionally, the user provides the security research and judgment problem text in the mode of direct input, and the user directly inputs the problem to the model in the question and answer box. For example, the user types the problem text in a text box, and then the model receives the problem and generates an answer, which is the most direct way to obtain the security research and judgment problem text.

[0129] Optionally, referring to Figure 6b In the scheme, first, the original problem text is obtained, that is, the security research and judgment problem text is obtained. Then, the large language model generates an answer. Subsequently, the large language model summarizes the demand according to the one question and one answer text, and performs processing on the target security data that matches the security research and judgment demand information, to obtain the processed security data. When the target security data is a security data report, the large language model processes the security data report by summarizing the demand, and finally generates the security research and judgment answer corresponding to the security research and judgment problem text.

[0130] For the obtained security research and judgment problem text, the large language model can understand the problem by analyzing features. For example, by analyzing vocabulary, syntax, and context to grasp the meaning of the problem. This usually involves natural language processing (NLP) techniques such as word segmentation, part-of-speech tagging, dependency parsing, etc. After understanding the problem, the model needs to search its internal knowledge base to find information related to the problem (in actual application, it can refer to finding a set of portals containing various security systems). The model generates an answer based on the information found. This process usually involves an algorithm called "decoding", in which the model selects the most appropriate answer from the possible answers. For some advanced models such as Transformer, they will use a series of attention mechanisms to focus on different parts of the input text and generate coherent answers. The process of generating a security research and judgment answer is a complex and accuracy-oriented process. By effectively performing the above steps, better use of data can be made to support security operations.

[0131] Optionally, generating a security research and judgment answer can be achieved by the following methods: Requirement understanding: Clearly define the requirements of the research and judgment, including which data is needed, how the data is displayed, the format and structure of the research and judgment, etc.

[0132] Data collection: After the requirements are clear, the model collects relevant data. These data can come from third-party integrated major security systems.

[0133] Data processing: The collected data often needs to be processed to meet the requirements of the research and judgment. This may include data cleaning (removing duplicate and incorrect data), data conversion (such as formatting dates, unifying units, etc.), data aggregation (such as sum, average, etc.) operations.

[0134] Data modeling: After processing the data, the model may perform data modeling according to the requirements, such as building prediction models, association rule models, etc. This step can help the model better understand the data and provide support for subsequent research and judgment generation.

[0135] Finally, the model generates a security research and judgment answer based on the requirements and processed data. Before that, the results of the model's demand analysis can be corrected to ensure the accuracy and reliability of the final security research and judgment answer generated by the model.

[0136] The scheme provided in the embodiments of the present application applies a semantic matching algorithm, which can be used to understand and compare the semantic similarity between text or information segments. The algorithm can effectively analyze the text content, extract key concepts and construct a semantic network, and then calculate the similarity score between different texts. A high score indicates that the two texts are very close in meaning. The scheme provided in the embodiments of the present application is easy to deploy and has high scalability, stability and security in a security research and judgment system, which is conducive to adapting to the growing data volume and technical requirements.

[0137] Since the multi-source security data can include different forms of data, the data security research and judgment system can integrate multiple data sources and service interfaces to facilitate interoperability with other business systems. In addition, at the security level, strict identity verification mechanisms and access control policies can be implemented to prevent unauthorized information leakage. The present scheme comprehensively applies semantic matching algorithms and robust system architecture design to effectively improve the effectiveness of data security research and judgment.

[0138] To solve the problems in the related art, the embodiments of the present application also provide a data security research and judgment apparatus 70, as shown in Figure 7 The apparatus 70 comprises: an acquisition module 71 configured to acquire a security research and judgment question text; an execution module 72 configured to perform text semantic analysis on the security research and judgment question text to obtain security research and judgment requirement information and data source information of a semantic representation of the security research and judgment question text; a determination module 73 configured to determine target security data matched with the data source information from a plurality of preset security data sources; a processing module 74 configured to perform processing matched with the security research and judgment requirement information on the target security data to obtain processed security data; a generation module 75 configured to generate a security research and judgment answer responding to the security research and judgment question text based on the processed security data.

[0139] The device provided in the embodiment of the present application firstly acquires a security research question text; then performs text semantic analysis on the security research question text to obtain security research demand information and data source information of semantic representation of the security research question text; next, determines target security data matched with the data source information from preset multi-source security data; subsequently, performs processing matched with the security research demand information on the target security data to obtain processed security data; and finally, generates a security research answer responding to the security research question text based on the processed security data. Through the scheme provided in the embodiment of the present application, the security research question text can be effectively analyzed in the security research scene, so as to determine the security research demand information and the data source information. In this way, the target security data can be effectively acquired and the processing matched with the security research demand information can be performed in the subsequent steps. Thus, the security data obtained through the processing can meet the semantic demand of the security research question text, and the security research answer can effectively respond to the security research question, thereby improving the effectiveness of data security research.

[0140] In the device provided in the embodiment of the present application, the above modules can also implement the method steps provided in the method embodiments. Alternatively, the device provided in the embodiment of the present application can also include other modules in addition to the above modules to implement the method steps provided in the method embodiments. The device provided in the embodiment of the present application can achieve the technical effects achievable by the method embodiments.

[0141] Preferably, the embodiment of the present application further provides an electronic device, including a processor, a memory, a computer program stored in the memory and executable on the processor, which implements each process of the above-mentioned data security research method embodiments and achieves the same technical effects when executed by the processor. To avoid repetition, details are not described here.

[0142] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement each process of the above-mentioned data security research method embodiments and achieve the same technical effects. To avoid repetition, details are not described here. The computer readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0143] The embodiment of the present application further provides a computer program product, which includes a non-transitory computer readable storage medium storing a computer program. The computer program is operable to cause a computer to perform some or all of the steps of the above-mentioned data security research method embodiments and achieve the same technical effects. To avoid repetition, details are not described here.

[0144] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, apparatus, or computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer readable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0145] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks.

[0146] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart illustrations and / or block diagrams block or blocks.

[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks.

[0148] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0149] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, non-volatile memory, such as read-only memory (ROM) or flash memory, etc. Memory is an example of computer readable media.

[0150] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0151] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0152] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0153] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. A data security research and judgment method, characterized in that, The method comprises the following steps: obtaining a security research and judgment question text; performing text semantic analysis on the security research and judgment question text to obtain security research and judgment requirement information and data source information of semantic representation of the security research and judgment question text; determining target security data matched with the data source information from preset multi-source security data; performing processing matched with the security research and judgment requirement information on the target security data to obtain processed security data; generating a security research and judgment answer responding to the security research and judgment question text based on the processed security data.

2. The method of claim 1, wherein, Before determining the target security data matched with the data source information from the preset multi-source security data, the method further comprises the following steps: obtaining security data to be processed from a plurality of data sources respectively; performing correlation analysis on the security data to be processed from different data sources to obtain correlation information between the security data; constructing the multi-source security data based on the correlation information and the security data to be processed from different data sources.

3. The method of claim 1, wherein, Before performing correlation analysis on the security data to be processed from different data sources to obtain correlation information between the security data, the method further comprises the following steps: performing anomaly detection on the security data to be processed from a plurality of data sources by an isolation forest algorithm; labeling security data detected as abnormal as an abnormal state; wherein, performing correlation analysis on the security data to be processed from different data sources to obtain correlation information between the security data comprises: performing correlation analysis on the security data to be processed from different data sources based on the security data labeled as the abnormal state to obtain correlation information between the security data.

4. The method of claim 1, wherein, Determining the target security data matched with the data source information from the preset multi-source security data comprises: performing vector similarity calculation on the preset multi-source security data and the security research and judgment requirement information to obtain first security data matched in vectors; performing semantic matching judgment on the preset multi-source security data and the security research and judgment requirement information to obtain second security data matched in semantics; determining the target security data based on the first security data and the second security data.

5. The method of claim 1, wherein, The security research and judgment question text comprises a plurality of texts based on time sequence; wherein, performing text semantic analysis on the security research and judgment question text to obtain security research and judgment requirement information and data source information of semantic representation of the security research and judgment question text comprises: performing keyword analysis on the security research and judgment question text to obtain keywords corresponding to the plurality of texts based on time sequence respectively; determining keyword similarity between texts based on the keywords corresponding to the plurality of texts respectively; performing semantic analysis on the plurality of texts respectively to determine security research and judgment requirement information probability and data source information probability corresponding to the plurality of texts respectively; determining security research and judgment requirement information and data source information of the security research and judgment question text based on the keyword similarity between texts and the security research and judgment requirement information probability and the data source information probability corresponding to the plurality of texts respectively.

6. The method of claim 1, wherein, The target security data comprises user local privacy data; wherein, performing processing matched with the security research and judgment requirement information on the target security data to obtain processed security data comprises: The local privacy data of the user is processed by a large language model deployed locally to the user to obtain processed security data.

7. A device for data security research and judgment, characterized in that, The method comprises the following steps: An acquisition module acquires security research and judgment question text. An execution module performs text semantic analysis on the security research and judgment question text to obtain security research and judgment requirement information and data source information of the security research and judgment question text semantic representation. A determination module determines target security data matched with the data source information from a plurality of preset security data sources. A processing module performs processing on the target security data matched with the security research and judgment requirement information to obtain processed security data. A generation module generates a security research and judgment answer responding to the security research and judgment question text based on the processed security data.

8. An electronic device, comprising: The method comprises the following steps: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is executed by the processor to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium and is executed by the processor to implement the steps of the method according to any one of claims 1 to 6.

10. A computer program product, characterised in that, The computer program product comprises a non-transitory computer readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the steps of the method according to any one of claims 1 to 6.