A method for generating a semantic association analysis model and a semantic association analysis method

By using a semantic association analysis model and deep neural networks to process and predict unstructured data, the problem of low efficiency in traditional case investigation has been solved, enabling efficient and accurate data analysis and insights, and adapting to complex and ever-changing data environments.

CN119227693BActive Publication Date: 2026-03-13SHANGHAI XINREN INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In traditional case investigations, manual analysis is inefficient and difficult to be comprehensive and accurate when faced with complex and diverse unstructured data. Traditional analysis methods lack flexibility and cannot quickly adapt to changing data and information, resulting in a time-consuming and slow-to-result process.

Method used

A semantic association analysis model is adopted, including modules for data acquisition, preprocessing, comprehensive semantic prediction and analysis. Deep neural networks are used for data processing and prediction, and a visualization module is combined to achieve accurate prediction and dynamic updates for various types of data.

Benefits of technology

It enables efficient and accurate analysis of unstructured data, improving the efficiency and accuracy of case investigation, and can deeply explore the intrinsic connections between data to provide users with comprehensive and in-depth insights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119227693B_ABST
    Figure CN119227693B_ABST
Patent Text Reader

Abstract

This invention discloses a method for generating a semantic association analysis model and a semantic association analysis method, comprising: acquiring historical target data and preprocessing the historical target data; establishing a comprehensive text semantic analysis model based on a deep neural network; training the comprehensive text semantic analysis model based on the preprocessed historical target data; inputting real-time target data and preprocessing the real-time target data, inputting the preprocessed real-time target data into the trained comprehensive text semantic analysis model to obtain the output result; and performing semantic association analysis on the output result. The comprehensive semantic prediction model achieves accurate prediction of various types of target data, and continuously optimizes model performance through a dynamic update mechanism to ensure the accuracy and timeliness of the prediction results. Its unique semantic association analysis module can deeply explore the inherent connections between different types of data such as text, images, audio, and video, providing users with more comprehensive and in-depth insights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text semantic association analysis technology, and in particular to a method for generating a semantic association analysis model and a semantic association analysis method. Background Technology

[0002] With the rapid development of technology and the comprehensive digital transformation of society, electronic data has become indispensable in criminal investigations. However, the rapid advancement of information technology has also brought about an extremely complex data environment, resulting in an explosive growth in the amount of data involved in cases. In this process, not only has the amount of data surged, but the types and formats of data have also become increasingly diversified. A large amount of unstructured data, such as audio and video, images, documents, and data from social media, mobile devices, and cloud services, poses unprecedented challenges to traditional criminal investigation data governance and analysis.

[0003] Faced with such a complex data environment, traditional analytical methods prove inadequate. Traditional methods typically rely on manual data screening, processing, and analysis, which is not only labor-intensive and inefficient but also easily constrained by human factors, leading to incomplete and inaccurate results. Furthermore, the manual data processing is often only suitable for structured data, such as tables and records, and has very limited capabilities for processing unstructured data such as audio, video, and images. In addition, when dealing with new and complex case scenarios, traditional analytical methods lack specificity and flexibility, failing to quickly adapt to and respond to changing data and information, resulting in time-consuming and slow-to-result processes. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0005] In view of the aforementioned existing problems, the present invention is proposed.

[0006] Therefore, the present invention provides a method for generating a semantic association analysis model and a semantic association analysis method, which can solve the problems mentioned in the background art.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0008] In a first aspect, the present invention provides a method for generating a semantic association analysis model, the semantic association analysis model comprising: a data acquisition module, a data preprocessing module, a comprehensive semantic prediction model module, a comprehensive semantic analysis module, and a visualization module;

[0009] The data acquisition module is used to acquire first target data and transmit the first target data to the data preprocessing module;

[0010] The data preprocessing module is used to perform a first classification on the first target data transmitted by the data acquisition module to obtain the second target data, and to perform a first preprocessing on the second target data to obtain the third target data, and at the same time transmit the third target data to the comprehensive semantic prediction model module;

[0011] The comprehensive semantic prediction model module is used to perform comprehensive semantic prediction on the third target data transmitted by the data preprocessing module through a comprehensive semantic prediction model obtained based on deep neural network pretraining, and transmit the prediction results of the comprehensive semantic prediction model to the comprehensive semantic analysis module.

[0012] The comprehensive semantic analysis module is used to analyze the prediction results transmitted by the comprehensive semantic prediction model module, and transmit the prediction results after the analysis is completed to the visualization module;

[0013] The visualization module is used to perform visualization processing on the prediction results after the analysis is completed, which are transmitted from the comprehensive semantic analysis module.

[0014] As a preferred embodiment of the semantic association analysis model generation method described in this invention, the comprehensive semantic prediction model module includes: a data processing and input unit, a comprehensive semantic prediction model unit, a comprehensive semantic prediction model update unit, and a data storage unit.

[0015] The data processing and input unit is used to perform a second preprocessing on the third target data transmitted by the data preprocessing module to obtain a fourth target data, and to perform a second classification on the fourth target data to obtain a fifth target data, while transmitting the fifth target data to the comprehensive semantic prediction model unit and the data storage unit.

[0016] The comprehensive semantic prediction model unit performs model prediction based on the transmitted fifth target data, obtains prediction results of different models, and transmits the prediction results to the comprehensive semantic analysis module and the data storage unit;

[0017] The data storage unit is used to store the fifth target data generated by the data processing and input unit and the prediction results generated by the comprehensive semantic prediction model unit for the fifth target data.

[0018] As a preferred embodiment of the semantic association analysis model generation method described in this invention, the comprehensive semantic prediction model update unit includes:

[0019] The integrated semantic prediction model update unit has a preset update cycle. When the update cycle is reached and the model needs to be updated, the first correct data stored in the data storage unit within that cycle is retrieved to update the model, and the updated model is transmitted to the integrated semantic prediction model unit to replace the original model.

[0020] The first correct data is the fifth target data generated by the data processing and input unit whose prediction result within the period is accurate, and the prediction result generated by the comprehensive semantic prediction model unit for the fifth target data.

[0021] As a preferred embodiment of the method for generating the semantic association analysis model described in this invention, the comprehensive semantic prediction model unit includes: a first text prediction model unit, a second text prediction model unit, and a third text prediction model unit;

[0022] The first text prediction model is used to predict the first text data in the fifth target data, the second text prediction model is used to predict the second text data in the fifth target data, and the third text prediction model is used to predict the third text data in the fifth target data.

[0023] As a preferred embodiment of the method for generating the semantic association analysis model described in this invention, the comprehensive semantic analysis module includes: a first analysis unit, a second analysis unit, a third analysis unit, and a fourth analysis unit;

[0024] The first analysis unit is used to receive the prediction result transmitted by the first text prediction model unit, and to perform first text analysis on the prediction result;

[0025] The second analysis unit is used to receive the prediction results transmitted by the second text prediction model unit, and to perform second text analysis on the prediction results;

[0026] The third analysis unit is used to receive the prediction results transmitted by the third text prediction model unit, and to perform third text analysis on the prediction results;

[0027] The fourth analysis unit is used to perform a first correlation analysis on the first text analysis result, the second text analysis result, and the third text analysis result, and transmit the first correlation analysis to the visualization module for visualization processing.

[0028] As a preferred embodiment of the method for generating the semantic association analysis model described in this invention, the data preprocessing module includes: a first classification unit, a first preprocessing unit, a first data output unit, and a first preprocessing adjustment unit;

[0029] The first classification unit is used to perform a first classification on the first target data according to a preset classification rule;

[0030] The first preprocessing unit is used to perform a first preprocessing on the second target data after the first classification, and generate the third target data;

[0031] The first data output unit transmits the preprocessed third target data to the comprehensive semantic prediction model module;

[0032] The first preprocessing adjustment unit is used to obtain the first erroneous data stored in the data storage unit in the comprehensive semantic prediction model module, and adjust the preprocessing parameters of the first preprocessing unit according to the first erroneous data;

[0033] The first erroneous data is the fifth target data generated by the data processing and input unit whose prediction result is incorrect within the period, and the prediction result generated by the comprehensive semantic prediction model unit for the fifth target data.

[0034] Secondly, the present invention provides a semantic association analysis method for generating a semantic association analysis model, comprising:

[0035] Acquire historical target data and preprocess the historical target data;

[0036] A comprehensive text semantic analysis model is established based on deep neural networks;

[0037] The comprehensive text semantic analysis model is trained based on the preprocessed historical target data;

[0038] Input real-time target data and preprocess the real-time target data. Then, input the preprocessed real-time target data into the trained comprehensive text semantic analysis model to obtain the output results.

[0039] Perform semantic association analysis on the output results.

[0040] As a preferred embodiment of the semantic association analysis method for generating the semantic association analysis model described in this invention, the comprehensive text semantic analysis model includes: a first text prediction model, a second text prediction model, and a third text prediction model.

[0041] The first text prediction model is used to predict text-type target data;

[0042] The second text prediction model is used to predict data where the target data is image-based;

[0043] The third text prediction model is used to predict target data that is audio or video.

[0044] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.

[0045] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0046] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention proposes a method for generating a semantic association analysis model and a semantic association analysis method. The method involves acquiring historical target data and preprocessing it; establishing a comprehensive text semantic analysis model based on a deep neural network; training the comprehensive text semantic analysis model based on the preprocessed historical target data; inputting real-time target data and preprocessing it, while simultaneously inputting the preprocessed real-time target data into the trained comprehensive text semantic analysis model to obtain the output result; and performing semantic association analysis on the output result. This semantic association analysis model and method not only achieve accurate prediction of various types of target data through a comprehensive semantic prediction model, but also continuously optimizes model performance through a dynamic update mechanism, ensuring the accuracy and timeliness of the prediction results. Furthermore, its unique semantic association analysis module can deeply explore the inherent connections between different types of data such as text, images, and audio / video, providing users with more comprehensive and in-depth insights. Attached Figure Description

[0047] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0048] Figure 1 This is a schematic diagram of the system structure of a semantic association analysis model generation method and a semantic association analysis method provided in an embodiment of the present invention;

[0049] Figure 2 A flowchart illustrating a method for generating a semantic association analysis model and a semantic association analysis method, provided as an embodiment of the present invention;

[0050] Figure 3 This is an internal structural diagram of a computer device for generating a semantic association analysis model and performing semantic association analysis, as provided in an embodiment of the present invention. Detailed Implementation

[0051] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0052] Example 1

[0053] Reference Figures 1-3 This is the first embodiment of the present invention, which provides a method for generating a semantic association analysis model and a semantic association analysis method, including:

[0054] Before detailing the embodiments of this application, some related concepts will be explained for clarity.

[0055] Semantic association analysis is a method of information processing that focuses on identifying, understanding, and extracting meaningful relationships between different datasets. These relationships can exist in various types of data, such as text, images, audio, and video. Through semantic association analysis, useful information can be extracted from unstructured data, revealing the inherent connections between them.

[0056] Deep neural networks: This is a subfield of machine learning that mimics the working principle of neurons in the human brain to construct multi-layered artificial neural networks. Deep neural networks handle complex learning tasks through multi-layered abstract representations. They are applicable to various fields such as image recognition, speech recognition, and natural language processing, and have powerful pattern recognition and self-learning capabilities.

[0057] Association analysis is a data mining technique that aims to discover interesting associations or frequently occurring patterns between variables in large datasets. For example, market basket analysis is an application of association analysis that helps retailers understand the relationships between customer buying behaviors—that is, what other products customers frequently buy after purchasing one item.

[0058] Long context learning: This term typically refers to the ability of an algorithm to take into account information over a long time span or paragraph when processing sequential data (such as text or time series). In natural language processing, long context learning means that the model is able to understand information across multiple sentences or paragraphs, which is crucial for tasks requiring contextual understanding (such as question answering systems and text summarization).

[0059] Natural Language Processing (NLP) is a branch of artificial intelligence that focuses on enabling computers to understand, interpret, and generate human language. NLP combines linguistic knowledge, computer science, and mathematical algorithms, allowing machines to process and understand large amounts of natural language data, thereby performing tasks such as automatic translation, sentiment analysis, and question-answering systems.

[0060] In related technologies, with the deep penetration and widespread application of the internet, cybercrime cases are showing a rapid growth trend. These cases typically possess the following characteristics: large scale, numerous individuals involved, and diverse types of evidence. These characteristics lead to a dramatic increase in the volume and complexity of data when processing and analyzing case data, requiring investigators to handle massive amounts, even tens of millions, of social media content and communication data. Traditional processing methods mainly rely on manual review, comparison, and analysis, which is not only time-consuming and labor-intensive but also difficult to guarantee accuracy and comprehensiveness. With the increasing volume and complexity of case data, traditional methods can no longer meet the needs of modern case investigation.

[0061] This application provides a solution to the problems mentioned above. The following will describe in detail how to generate the semantic association analysis model by combining multiple embodiments.

[0062] Figure 1 A method for generating a semantic association analysis model and a schematic diagram of the system structure of the semantic association analysis method are shown, including: a data acquisition module 100, a data preprocessing module 200, a comprehensive semantic prediction model module 300, a comprehensive semantic analysis module 400, and a visualization module 500;

[0063] In this embodiment of the application, the data acquisition module 100 is used to acquire the first target data and transmit the first target data to the data preprocessing module 200;

[0064] The data preprocessing module 200 is used to perform a first classification on the first target data transmitted from the data acquisition module 100 to obtain the second target data, and to perform a first preprocessing on the second target data to obtain the third target data. At the same time, the third target data is transmitted to the comprehensive semantic prediction model module 300.

[0065] The comprehensive semantic prediction model module 300 is used to perform comprehensive semantic prediction on the third target data transmitted from the data preprocessing module 200 through a comprehensive semantic prediction model obtained based on deep neural network pre-training, and transmits the prediction results of the comprehensive semantic prediction model to the comprehensive semantic analysis module 400.

[0066] The comprehensive semantic analysis module 400 is used to analyze the prediction results transmitted from the comprehensive semantic prediction model module 300, and transmit the prediction results after the analysis to the visualization module 500.

[0067] The visualization module 500 is used to perform visualization processing on the prediction results after the analysis is completed, which are transmitted from the comprehensive semantic analysis module 400.

[0068] It should be noted that if a semantic association analysis model of the type required by the client needs to be generated, the data acquisition module 100 only needs to import the corresponding type of data source (i.e., the first target data), and the input and output of the comprehensive semantic prediction model module 300, which is pre-trained based on a deep neural network, need to be changed. For example, when generating a semantic association analysis model for case investigation, case investigation-related data sources can be imported, such as social media content, communication records, and information on suspects. Simultaneously, the comprehensive semantic prediction model in the comprehensive semantic prediction model module 300 needs to be adjusted to enable effective semantic prediction and analysis for these specific types of data. For example, the input of the comprehensive semantic prediction model module 300 can be changed to text data related to case investigation, such as the suspect's communication records and social media content, while the model's output can be adjusted to the semantic association analysis results of the case-related content. Such adjustments allow the model to focus more on the specific needs of the case investigation field, improving the accuracy and efficiency of the analysis.

[0069] In an optional embodiment, the data acquisition module 100 can incorporate a flexible data source interface to support the capture of information from diverse data sources, including but not limited to social media platforms, instant messaging applications, email systems, and public safety databases. These data sources cover a wide range of information that may be involved in the investigation process, ensuring the comprehensiveness and timeliness of the data.

[0070] In this embodiment of the application, the data preprocessing module 200 includes: a first classification unit 201, a first preprocessing unit 202, a first data output unit 203, and a first preprocessing adjustment unit 204;

[0071] The first classification unit 201 is used to perform a first classification on the first target data according to a preset classification rule;

[0072] The first preprocessing unit 202 is used to perform first preprocessing on the second target data after the first classification, and generate third target data;

[0073] The first data output unit 203 transmits the preprocessed third target data to the comprehensive semantic prediction model module 300;

[0074] The first preprocessing adjustment unit 204 is used to obtain the first error data stored in the data storage unit 304 in the comprehensive semantic prediction model module 300, and adjust the preprocessing parameters of the first preprocessing unit 202 according to the first error data.

[0075] The first error data is the fifth target data generated by the data processing and input unit 301 with an incorrect prediction result within a cycle, and the prediction result generated by the comprehensive semantic prediction model unit 302 for the fifth target data.

[0076] It should be noted that the first classification divides the first target data into three types: text data, picture data, and audio-video data.

[0077] In an optional embodiment, the first preprocessing includes, but is not limited to, processing text data such as removing stop words, word segmentation,词性标注, named entity recognition, etc., to improve the accuracy of subsequent semantic prediction; processing picture data such as image recognition, feature extraction, etc., so as to convert the key information in the image into text information that can be processed by the model; processing audio-video data such as speech recognition, transcription, keyword extraction, etc., so as to extract useful text information from the audio-video.

[0078] In the embodiment of the present application, when a semantic association analysis model of a case investigation type needs to be generated, the process of the first preprocessing for text data is as follows: First, remove common words in the text, such as "的", "了", "是", which have little impact on understanding the meaning of the text. Second, split the continuous text string into individual words or phrases for subsequent processing. Third, mark the词性 of each word, such as noun (N), verb (V), adjective (J), etc. Finally, identify proper nouns in the text, such as names of people, places, organizations, etc., and classify them.

[0079] Exemplarily, suppose there is a text: "张三在M市银行盗窃案中被捕。"

[0080] First, remove "的", "在", "中", "被".

[0081] Second, split the text into "张三", "M市", "银行", "盗窃案", "被捕".

[0082] Third, they are respectively "name of person", "place name", "noun", "event", "verb".

[0083] Finally, identify that "张三" is the name of a person and "M市" is the name of a place.

[0084] In the embodiment of the present application, when a semantic association analysis model of a case investigation type needs to be generated, the process of the first preprocessing for picture data is as follows: First, use computer vision technology to identify objects or scenes in the picture. Second, extract key features from the picture, such as edges, color distribution, texture, etc. Finally, convert the information in the image into a text description for subsequent processing.

[0085] Exemplarily, suppose there is a picture containing a car and a license plate number. It should be noted that the Chinese word "词性标注" in the original text seems to be an incomplete or incorrect expression. It is guessed that it may be "词性标注" (POS tagging). The above translation is based on this guess. If there is an error, please correct it according to the correct information.

[0086] First, identify the "car" and "license plate number".

[0087] Next, extract features such as the car's color and model.

[0088] Finally, the image was described as "a red BMW with license plate number XX".

[0089] In this embodiment of the application, when it is necessary to generate a semantic association analysis model of the case investigation type, the first preprocessing of the audio and video data is as follows: First, the speech in the audio is converted into text. Second, the recorded content is converted into written form. Finally, key information is extracted from the transcribed text.

[0090] For example, suppose there is a recording: "The suspect admits that he was at the scene around 10 p.m. on the night of the incident."

[0091] Speech recognition: Converts audio recordings into text.

[0092] Transcription: The obtained text reads, "The suspect admitted that he was at the scene around 10 p.m. on the night of the incident."

[0093] Keyword extraction: Extracted keywords "suspect", "around 10 p.m. on the night of the incident", and "appeared at the scene".

[0094] It should be noted that the first preprocessing described above can perfectly preprocess three different types of data: text data, image data, and audio / video data, providing a solid foundation for subsequent semantic prediction analysis.

[0095] It should also be noted that the first type of error data refers to predictions that are incorrect within a given period. The judgment of whether a prediction is correct or incorrect can be made manually or by setting a preset threshold. Manual judgment simply requires randomly checking whether the generated prediction is satisfactory. If there are unsatisfactory parts, the model prediction is directly marked as incorrect or inaccurate. Although manual judgment is possible at this stage, it requires less time and effort, almost negligible compared to traditional manual comparison processes, thus avoiding excessive manpower consumption. Setting a preset threshold involves setting a threshold based on the deviation between the model's prediction and the actual result. When the deviation exceeds this threshold, the model's prediction is considered incorrect. This method is relatively objective, but requires setting an appropriate threshold in advance to ensure the accuracy of the judgment.

[0096] It should be noted that the generation method of the above model and the operation process of the system both demonstrate a high degree of modularity and flexibility. This design not only facilitates system maintenance and upgrades but also enables the system to be quickly adapted to different application scenarios and needs. For example, when dealing with emergencies or analyzing specific industries, the system can be quickly reconstructed to meet new analytical requirements simply by adjusting the corresponding module parameters or replacing specific modules.

[0097] In this embodiment of the application, the comprehensive semantic prediction model module 300 includes: a data processing and input unit 301, a comprehensive semantic prediction model unit 302, a comprehensive semantic prediction model update unit 303, and a data storage unit 304.

[0098] The data processing and input unit 301 is used to perform a second preprocessing on the third target data transmitted by the data preprocessing module 200 to obtain the fourth target data, and to perform a second classification on the fourth target data to obtain the fifth target data. At the same time, the fifth target data is transmitted to the comprehensive semantic prediction model unit 302 and the data storage unit 304.

[0099] The comprehensive semantic prediction model unit 302 performs model prediction based on the transmitted fifth target data, obtains the prediction results of different models, and transmits the prediction results to the comprehensive semantic analysis module 400 and the data storage unit 304.

[0100] The data storage unit 304 is used to store the fifth target data generated by the data processing and input unit 301 and the prediction results generated by the comprehensive semantic prediction model unit 302 for the fifth target data.

[0101] In an optional embodiment, the second preprocessing includes, but is not limited to, data cleaning, format unification, and standardization, which aim to further improve the accuracy and consistency of the data and provide more reliable data input for the comprehensive semantic prediction model.

[0102] In this embodiment, the second preprocessing is to ensure that the data input to the model better matches the model's expectations and requirements, reducing model errors and thus improving the accuracy of model predictions. Specifically, data cleaning removes missing, abnormal, or duplicate data to ensure the quality of the input data; format unification converts all data into a model-recognizable format to avoid errors caused by inconsistent formats; and standardization scales the data according to a certain ratio, allowing data of different magnitudes to be compared and calculated fairly in the model.

[0103] It should be noted that the second classification is a specific classification operation. For example, for the investigation of specific types of cases, the second classification can divide the data after the second preprocessing into multiple subcategories based on the nature of the case, the fields involved, or the purpose of the analysis. For instance, in economic crime cases, the data can be divided into multiple subcategories such as financial transaction records, communication records, and electronic evidence; in criminal cases, it can be divided into subcategories such as crime scene investigation records, suspect confessions, and witness testimonies. This classification not only helps to analyze cases more precisely but also improves the relevance and accuracy of the model.

[0104] In this embodiment, during the second classification, text classification algorithms from Natural Language Processing (NLP) techniques, such as Naive Bayes, Support Vector Machine (SVM), and deep learning models, are utilized. These algorithms can automatically extract features from text data and assign texts to predefined categories based on these features. Furthermore, combining this with a specific knowledge base and rule set for case investigation can further improve the accuracy and efficiency of the classification.

[0105] For example, suppose we are processing a financial fraud case involving a large number of bank transaction and communication records. First, these records require a second preprocessing step, including data cleaning, format standardization, and data cleansing. During data cleaning, invalid, duplicate, or anomalous transaction records, such as those with incorrect amounts or timestamps, are removed. Simultaneously, communication records are filtered to remove irrelevant conversations and noise. In the format standardization stage, all transaction and communication records are converted to a standardized format, such as JSON or XML, for subsequent processing. In the standardization stage, numerical data such as transaction amounts and timestamps are scaled to ensure they are compared on the same order of magnitude.

[0106] Next, we proceed with the second classification. Based on the characteristics of financial fraud cases, transaction records can be divided into two subcategories: normal transactions and suspicious transactions. Normal transactions typically have reasonable amounts, times, and frequencies, while suspicious transactions may exhibit unusual characteristics, such as large transactions, nighttime transactions, or frequent fund transfers. Simultaneously, communication records can be divided into critical communications and non-critical communications. Critical communications may contain clues or evidence related to fraudulent activities, such as secret conversations or fraudulent instructions between suspects.

[0107] To perform the second classification, text classification algorithms from NLP techniques can be utilized. First, key features, such as transaction amount, time, location, and conversation content, are extracted from transaction and communication records. Then, these features are used to train a classification model that automatically assigns new transaction and communication records to the appropriate subcategories. During training, a specific knowledge base and rule set for financial fraud cases can be incorporated to improve the accuracy and efficiency of classification.

[0108] Once the second classification is complete, the classified data can be input into the comprehensive semantic prediction model for predictive analysis. The comprehensive semantic prediction model generates prediction results from different models based on the input data and a predefined set of rules. These prediction results may include information such as the probability of fraudulent behavior, the suspect's identity, and the methods of fraud. These prediction results can then be transmitted to the comprehensive semantic analysis module for further analysis and processing to support case investigation and detection.

[0109] In this embodiment of the application, the comprehensive semantic prediction model update unit 303 includes:

[0110] The integrated semantic prediction model update unit 303 has a preset update cycle. When the update cycle is reached and the model needs to be updated, it retrieves the first correct data stored in the data storage unit 304 within that cycle to update the model, and transmits the updated model to the integrated semantic prediction model unit 302 to replace the original model.

[0111] The first correct data is the fifth target data generated by the data processing and input unit 301, which is the accurate prediction result within the period, and the prediction result generated by the comprehensive semantic prediction model unit 302 for the fifth target data.

[0112] It should be noted that, due to the rapid development of technology and the comprehensive digital transformation of society, the amount of information data has exploded. Moreover, when it comes to case investigation, criminals come up with all sorts of new methods, so the model also needs to be updated in real time. Therefore, a comprehensive semantic prediction model update unit 303 was designed to achieve the periodic update of the model by setting a preset update cycle.

[0113] In this embodiment of the application, the comprehensive semantic prediction model unit 302 includes: a first text prediction model unit 302a, a second text prediction model unit 302b, and a third text prediction model unit 302c;

[0114] The first text prediction model unit 302a is used to predict the first text data in the fifth target data, the second text prediction model unit 302b is used to predict the second text data in the fifth target data, and the third text prediction model unit 302c is used to predict the third text data in the fifth target data.

[0115] It should be noted that the first text data is the corresponding data obtained from the plain text data in the first target data in the fifth target data. That is, the data in the first target data is itself plain text data, which is then processed by the first classification, the first preprocessing, the second preprocessing, and the second classification to obtain the fifth target data.

[0116] It should be noted that the second text data is the corresponding data obtained from the image data in the first target data in the fifth target data. That is, the data in the first target data that is itself image data is the fifth target data obtained after the first classification, the first preprocessing, the second preprocessing, and the second classification.

[0117] It should be noted that the third text data is the corresponding data obtained from the audio and video data in the first target data in the fifth target data. That is, the data in the first target data that is itself audio and video data is the fifth target data obtained after the first classification, the first preprocessing, the second preprocessing, and the second classification.

[0118] It should also be noted that because the accuracy rates of plain text data, image data, and audio / video data after being converted into text information are different, this application designs three text prediction models to achieve accurate recognition.

[0119] In the embodiments of this application, the prediction models in the first text prediction model unit 302a, the second text prediction model unit 302b, and the third text prediction model unit 302c are all designed using deep learning models based on the Transformer architecture. These models, such as GPT (Generative Pre-trained Transformer) or BERT (Bidirectional Encoder Representation Transformer), are used to capture the semantics and contextual relationships of the text. Through extensive training on large-scale case-related data, the model can generate high-quality embedding vectors or directly output prediction results, improving its ability to understand complex language (text). Furthermore, the designed prediction model can handle long sequences of text, thereby understanding long-distance dependencies between events and solving the information loss problem that occurs in traditional RNN or LSTM models when processing long texts.

[0120] For example, the prediction model in the first text prediction model unit 302a can be designed with the following architecture:

[0121] Input layer: The text data embedding layer maps individual words of the fifth target data and the first text data into vectors of fixed dimensions.

[0122] Position encoding: Position encoding is added to preserve the position information of the input sequence.

[0123] Multi-Head Self Attention (MHA): 6 layers, 8 heads per layer, used to capture long-distance dependencies in text.

[0124] Feed-Forward Network (FFN): Each layer contains two fully connected layers, connected by an activation function (such as ReLU).

[0125] Residual connections and layer normalization: Residual connections are added after MHA and FFN, and layer normalization is added before and after each sub-layer to stabilize the training process.

[0126] It should be noted that the connection method of this architecture is as follows:

[0127] Step 1: Input text -> Embedding Layer -> Positional Encoding;

[0128] Step 2: Position encoding -> Multi-head self-attention layer (MHALayer1) -> Residual connection -> Layer normalization -> Feedforward neural network (FFNLayer1) -> Residual connection -> Layer normalization;

[0129] Step 3: Repeat step 2 6 times;

[0130] Step 4: The Output Layer is used to generate the corresponding prediction results.

[0131] The prediction model in the second text prediction model unit 302b can be designed with the following architecture:

[0132] Input layer: An embedding layer for image data, which maps the fifth target data and the second text data into a text representation.

[0133] Multi-head self-attention layer: 4 layers, 8 heads per layer, used to process image-to-text conversion.

[0134] Feedforward neural networks: Each layer contains two fully connected layers, connected by an activation function (such as ReLU).

[0135] Residual connections and layer normalization: Residual connections are added after MHA and FFN, and layer normalization is added before and after each sub-layer to stabilize the training process.

[0136] It should be noted that the connection method of this architecture is as follows:

[0137] Step 1: Input image data, processed text data -> Embedding layer -> Multi-head self-attention layer (MHALayer1) -> Residual connection -> Layer normalization -> Feedforward neural network (FFNLayer1) -> Residual connection -> Layer normalization;

[0138] Step 2: Repeat Step 1 4 times;

[0139] Step 3: The Output Layer is used to generate the prediction results.

[0140] The prediction model in the third text prediction model unit 302c can be designed with the following architecture:

[0141] Input layer: An embedding layer for audio and video data, which maps the fifth target data and the third text data into a text representation.

[0142] Multi-head self-attention layer: 8 layers, 8 heads per layer, used to process audio / video to text conversion.

[0143] Feedforward neural networks: Each layer contains two fully connected layers, connected by an activation function (such as ReLU).

[0144] Residual connections and layer normalization: Residual connections are added after MHA and FFN, and layer normalization is added before and after each sub-layer to stabilize the training process.

[0145] It should be noted that the connection method of this architecture is as follows:

[0146] Step 1: Input audio and video data, processed text data -> Embedding layer -> Multi-head self-attention layer (MHALayer1) -> Residual connection -> Layer normalization -> Feedforward neural network (FFNLayer1) -> Residual connection -> Layer normalization;

[0147] Step 2: Repeat Step 1 8 times;

[0148] Step 3: The Output Layer is used to generate the prediction results.

[0149] It should be noted that the above example is only one possibility of the model framework. Relevant technicians may modify or redesign the framework according to specific circumstances or needs. However, regardless of the design, if the design process of this application is used, it should be within the scope of protection of this application.

[0150] In the embodiments of this application, the output of the model can include whether a person or account in the input first target data is involved in a case, whether there is a relationship, and whether there is suspicious behavior, that is, the probability of involvement in a case, the probability of a relationship, and the probability of suspicious behavior.

[0151] In this embodiment of the application, the comprehensive semantic analysis module 400 includes: a first analysis unit 401, a second analysis unit 402, a third analysis unit 403, and a fourth analysis unit 404;

[0152] The first analysis unit 401 is used to receive the prediction result transmitted by the first text prediction model unit 302a, and to perform first text analysis on the prediction result;

[0153] It should be noted that the first text analysis uses natural language processing (NLP) techniques, such as named entity recognition (NER), to specifically identify key figures, locations, times, and other elements in the case; at the same time, sentiment analysis techniques are used to assess the sentiment tendency in the text to help determine the possible trend of the nature of the case or the judgment result (i.e., the text involved in the case is analyzed and extracted according to the output probability of the model).

[0154] In this embodiment of the application, the second analysis unit 402 is used to receive the prediction result transmitted by the second text prediction model unit 302b, and perform second text analysis on the prediction result;

[0155] It should be noted that the second text analysis is responsible for processing the image text prediction results output by the second text prediction model unit 302b. Since image data often contains rich visual information and implicit contextual relationships, the second analysis unit needs to pay special attention to the interrelationships between key scenes, objects and their converted text in the image when parsing the image text (i.e., to analyze and extract the image data involved in the case based on the output probability of the model).

[0156] In this embodiment of the application, the third analysis unit 403 is used to receive the prediction result transmitted by the third text prediction model unit 302c, and perform third text analysis on the prediction result;

[0157] It should be noted that the third text analysis can identify and understand the speech content and speech cues in audio and video (currently, this application only considers the processing of the text model, converting audio and video or images into text information for processing). Since the accuracy of audio and video data after being converted into text data is low, the third text analysis is based on the lowest weight (i.e., the audio and video involved in the case are analyzed and extracted according to the output probability of the model).

[0158] In this embodiment of the application, the fourth analysis unit 404 is used to perform a first correlation analysis on the first text analysis result, the second text analysis result and the third text analysis result, and transmit the first correlation analysis to the visualization module 500 for visualization processing.

[0159] For example, suppose a social media account posts the message: "The suspect admits he was at the scene around 10 PM on the night of the incident," accompanied by a photo and a video. The first analysis unit 401 will then perform detailed natural language processing on this text. Through Named Entity Recognition (NER), the system can accurately identify the key figures in the text: "suspect," the key time "around 10 PM on the night of the incident," and the key location "the scene." This crucial information is essential for subsequent case analysis and prediction of the judgment. Simultaneously, sentiment analysis technology will be used to assess the sentiment tendency in the text. Although the text content is relatively objective in this scenario, and the sentiment tendency may not be obvious, this step still helps to comprehensively understand the background and possible course of the case.

[0160] Next, the second analysis unit 402 will take over processing the image-text prediction results of the accompanying images. Since image data often contains rich visual information and implicit contextual relationships, the second analysis unit needs to meticulously analyze the key scenes and objects in the images, as well as their relationship with the text. For example, the image may show the suspect's facial features, clothing, or the layout of the scene; this information can corroborate the text description, enhancing the credibility of the case analysis.

[0161] The third analysis unit 403 is responsible for processing the text information converted from audio and video data. Although the accuracy of converting audio and video data into text may not be as high as processing text data directly, the third analysis unit will still strive to identify and understand the speech content and verbal cues within it. These can provide valuable references for case analysis. Of course, in terms of weight allocation, the third analysis unit will give relatively low weights to audio and video data based on the model's output probability to ensure the accuracy and impartiality of the overall analysis.

[0162] Finally, the fourth analysis unit 404 will perform a first correlation analysis by integrating the results of the first, second, and third text analyses. In this step, the system will delve into the inherent connections and potential patterns between the various analysis units, such as the consistency between text descriptions and image content, and the complementarity between audio / video information and text descriptions. Through comprehensive correlation analysis, the system can generate a more complete and accurate case analysis report. This report includes the basic facts of the case, key evidence, and suspect information, as well as the suspect's current movements, location, possible appearance, attire, and the probability of identifying the suspect as such.

[0163] Ultimately, this case report, which integrates multi-source data and multi-dimensional analysis, will be transmitted to the visualization module 500 for visualization processing. Through intuitive and easy-to-understand charts and images, key information and analysis results of the case are presented, allowing users to more easily understand and grasp the overall picture and details of the case. Simultaneously, the visualization module also supports interactive operation and data filtering functions, allowing users to further explore and uncover hidden information and value within the case data according to their needs and interests.

[0164] This embodiment also provides a semantic association analysis method for a semantic association analysis model, including:

[0165] S101, acquire historical target data and preprocess the historical target data;

[0166] S102, A comprehensive text semantic analysis model is established based on deep neural networks;

[0167] S103, Train a comprehensive text semantic analysis model based on the preprocessed historical target data;

[0168] S104, Input real-time target data and preprocess the real-time target data. Then, input the preprocessed real-time target data into the trained comprehensive text semantic analysis model to obtain the output result.

[0169] S105, Perform semantic association analysis on the output results.

[0170] In this embodiment of the application, the comprehensive text semantic analysis model includes: a first text prediction model, a second text prediction model, and a third text prediction model;

[0171] The first text prediction model is used to predict data where the target data is text.

[0172] The second text prediction model is used to predict data where the target data is image-based.

[0173] The third text prediction model is used to predict target data that is audio or video.

[0174] The above-mentioned unit modules can be embedded in the processor of the computer device in hardware form or independent of it, or they can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.

[0175] This embodiment also provides a computer device, which may be a terminal, and its internal structure diagram may be as follows. Figure 3As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a semantic association analysis method for a semantic association analysis model. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0176] This embodiment also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, it performs the following steps:

[0177] Acquire historical target data and preprocess it;

[0178] A comprehensive text semantic analysis model is established based on deep neural networks;

[0179] A comprehensive text semantic analysis model is trained based on preprocessed historical target data;

[0180] Input real-time target data and preprocess the real-time target data. Then, input the preprocessed real-time target data into the trained comprehensive text semantic analysis model to obtain the output results.

[0181] Perform semantic association analysis on the output results.

[0182] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

[0183] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0184] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0185] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0186] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0187] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0188] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for generating a semantic association analysis model, characterized in that, The semantic association analysis model includes: The data acquisition module (100) is used to acquire first target data and transmit the first target data to the data preprocessing module (200). The first target data is case investigation data. The data preprocessing module (200) is used to perform a first classification on the first target data transmitted by the data acquisition module (100) to obtain the second target data, and to perform a first preprocessing on the second target data to obtain the third target data. At the same time, the third target data is transmitted to the comprehensive semantic prediction model module (300). The comprehensive semantic prediction model module (300) is used to perform comprehensive semantic prediction on the third target data transmitted by the data preprocessing module (200) through a comprehensive semantic prediction model obtained by pre-training based on a deep neural network, and transmits the prediction results of the comprehensive semantic prediction model to the comprehensive semantic analysis module (400); The comprehensive semantic prediction model module (300) includes a comprehensive semantic prediction model unit (302), which includes a first text prediction model unit (302a), a second text prediction model unit (302b), and a third text prediction model unit (302c). The first text prediction model unit (302a) contains a first text prediction model framework for text data, which maps the single words of the fifth target data, the first text data, into fixed-dimensional vectors as the input layer. Its framework contains a two-level architecture consisting of six position-encoding layers, multi-head self-attention layers, residual connection layers, normalization layers, feedforward neural network layers, residual connection layers, and normalization layers connected in sequence. The second text prediction model unit (302b) contains a second text prediction model framework for image data, which maps the fifth target data, the second text data, into text representations as the input layer. Its framework contains four two-level architectures consisting of four embedding layers, multi-head self-attention layers, residual connection layers, normalization layers, feedforward neural network layers, residual connection layers, and normalization layers connected in sequence. The third text prediction model unit (302c) contains a third text prediction model framework for audio and video data, which maps the fifth target data, the third text data, into text representations as the input layer. Its framework contains eight two-level architectures consisting of eight embedding layers, multi-head self-attention layers, residual connection layers, normalization layers, feedforward neural network layers, residual connection layers, and normalization layers connected in sequence. The comprehensive semantic analysis module (400) is used to analyze the prediction results transmitted by the comprehensive semantic prediction model module (300), and transmit the prediction results after the analysis to the visualization module (500); The visualization module (500) is used to visualize the prediction results after the analysis is completed, which are transmitted from the comprehensive semantic analysis module (400).

2. The method for generating the semantic association analysis model as described in claim 1, characterized in that, The comprehensive semantic prediction model module (300) further includes: a data processing and input unit (301), a comprehensive semantic prediction model update unit (303), and a data storage unit (304). The data processing and input unit (301) is used to perform a second preprocessing on the third target data transmitted by the data preprocessing module (200) to obtain a fourth target data, and to perform a second classification on the fourth target data to obtain a fifth target data, while transmitting the fifth target data to the comprehensive semantic prediction model unit (302) and the data storage unit (304); The comprehensive semantic prediction model unit (302) performs model prediction based on the transmitted fifth target data, obtains prediction results of different models, and transmits the prediction results to the comprehensive semantic analysis module (400) and the data storage unit (304); The data storage unit (304) is used to store the fifth target data generated by the data processing and input unit (301) and the prediction results generated by the comprehensive semantic prediction model unit (302) for the fifth target data.

3. The method for generating the semantic association analysis model as described in claim 2, characterized in that, The comprehensive semantic prediction model update unit (303) includes: The comprehensive semantic prediction model update unit (303) has a preset update cycle. When the update cycle is reached and the model needs to be updated, the first correct data stored in the data storage unit (304) within the cycle is retrieved to update the model, and the updated model is transmitted to the comprehensive semantic prediction model unit (302) to replace the original model. The first correct data is the fifth target data generated by the data processing and input unit (301) whose prediction result within the period is accurate, and the prediction result generated by the comprehensive semantic prediction model unit (302) for the fifth target data.

4. The method for generating the semantic association analysis model as described in claim 3, characterized in that, The comprehensive semantic prediction model unit (302) includes: a first text prediction model unit (302a), a second text prediction model unit (302b), and a third text prediction model unit (302c). The first text prediction model unit (302a) is used to predict the first text data in the fifth target data, the second text prediction model unit (302b) is used to predict the second text data in the fifth target data, and the third text prediction model unit (302c) is used to predict the third text data in the fifth target data.

5. The method for generating the semantic association analysis model as described in claim 4, characterized in that, The comprehensive semantic analysis module (400) includes: a first analysis unit (401), a second analysis unit (402), a third analysis unit (403), and a fourth analysis unit (404). The first analysis unit (401) is used to receive the prediction result transmitted by the first text prediction model unit (302a) and perform first text analysis on the prediction result; The second analysis unit (402) is used to receive the prediction result transmitted by the second text prediction model unit (302b) and perform second text analysis on the prediction result; The third analysis unit (403) is used to receive the prediction result transmitted by the third text prediction model unit (302c) and perform third text analysis on the prediction result; The fourth analysis unit (404) is used to perform a first correlation analysis on the first text analysis result, the second text analysis result and the third text analysis result, and transmit the first correlation analysis to the visualization module (500) for visualization processing.

6. The method for generating the semantic association analysis model as described in claim 5, characterized in that, The data preprocessing module (200) includes: a first classification unit (201), a first preprocessing unit (202), a first data output unit (203), and a first preprocessing adjustment unit (204); The first classification unit (201) is used to perform a first classification on the first target data according to a preset classification rule; The first preprocessing unit (202) is used to perform a first preprocessing on the second target data after the first classification, and generate the third target data; The first data output unit (203) transmits the preprocessed third target data to the comprehensive semantic prediction model module (300). The first preprocessing adjustment unit (204) is used to obtain the first error data stored in the data storage unit (304) in the comprehensive semantic prediction model module (300), and adjust the preprocessing parameters of the first preprocessing unit (202) according to the first error data; The first erroneous data is the fifth target data generated by the data processing and input unit (301) whose prediction result is wrong within the period, and the prediction result generated by the comprehensive semantic prediction model unit (302) for the fifth target data.

7. A semantic association analysis method applied to the semantic association analysis model of claim 1, characterized in that, The method includes: Acquire historical target data and preprocess the historical target data; A comprehensive text semantic analysis model is established based on deep neural networks; The comprehensive text semantic analysis model is trained based on the preprocessed historical target data; Input real-time target data and preprocess the real-time target data. Then, input the preprocessed real-time target data into the trained comprehensive text semantic analysis model to obtain the output results. Perform semantic association analysis on the output results.

8. The semantic association analysis method of the semantic association analysis model as described in claim 7, characterized in that, The comprehensive text semantic analysis model includes: a first text prediction model, a second text prediction model, and a third text prediction model; The first text prediction model is used to predict text-type target data; The second text prediction model is used to predict data where the target data is image-based; The third text prediction model is used to predict target data that is audio or video.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 7 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 7 to 8.

Citation Information

Patent Citations

  • Anti-smuggling case information extraction method based on big data

    CN111476027A

  • Method and system for automatically generating subtitles in short video

    CN113159034A

  • Semantic analysis-based inspection and detection service method

    CN114925697A