Target object recognition method and device, electronic equipment and storage medium

Through large language models and heterogeneous graph construction technology, combined with attention mechanism and cross-modal fusion, the lag problem of social network dynamic analysis in existing technologies is solved, and the rapid and accurate identification of potential high-risk targets is achieved, supporting public safety and platform governance.

CN120597264APending Publication Date: 2025-09-05BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510384236.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies rely on static analysis models when identifying potential high-risk targets on social networks. These models cannot effectively reflect the dynamic characteristics of user interactions and posting interests over time, resulting in lags and misjudgments in the identification process.

Method used

A large language model is used for named entity recognition and topic extraction, and a sentence embedding representation model is used for normalization. Heterogeneous graph data is constructed, and behavioral evolution analysis is performed in the structural and temporal dimensions through the attention mechanism. Dynamic fusion is combined with the cross-modal attention mechanism, and finally risk assessment and ranking are performed through the recommendation model.

Benefits of technology

It has achieved rapid and accurate identification of potential high-risk targets in a large-scale dynamic social network environment, providing strong technical support for public security and platform governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597264A_ABST
    Figure CN120597264A_ABST
Patent Text Reader

Abstract

The invention provides a target object recognition method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining initial text data; wherein the initial text data comprises text data published by a plurality of to-be-identified objects; performing entity recognition and subject extraction on the initial text data by utilizing a large language model to obtain first intermediate data; inputting the first intermediate data into a sentence embedding representation model, and performing normalization processing on the first intermediate data by using the sentence embedding representation model to obtain second intermediate data; according to a set graph network construction rule, performing heterogeneous graph construction on the second intermediate data to generate heterogeneous graph data; performing behavior evolution analysis on the heterogeneous graph data in a structure dimension and a time dimension through an attention mechanism, and performing dynamic fusion on analysis results by using a cross-modal attention mechanism to obtain third intermediate data; and processing the third intermediate data through a recommendation model, screening the sorting result according to a preset screening strategy, and determining an identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic computers and Internet technologies, and in particular to a target object recognition method, device, electronic device, and storage medium. Background Art

[0002] With the rapid adoption of the internet and social media, the number of users worldwide has grown exponentially, forming a vast and complex information ecosystem. This has not only significantly improved the efficiency of information acquisition and dissemination, driving innovation in social civilization and economic models, but has also created security risks, with malicious users and criminal groups exploiting online anonymity to orchestrate illegal activities, posing a threat to online order, social stability, and public safety.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] In view of this, the present application proposes a target object recognition method, device, electronic device and storage medium to solve or partially solve the above problems.

[0005] Based on the above objectives, this application provides a target object recognition method, including:

[0006] Acquire initial text data; wherein the initial text data includes text data published by multiple objects to be identified;

[0007] Using a pre-trained large language model to perform entity recognition and topic extraction on the initial text data to obtain first intermediate data;

[0008] Inputting the first intermediate data into a pre-trained sentence embedding representation model, and normalizing the first intermediate data using the sentence embedding representation model to obtain second intermediate data;

[0009] According to the set graph network construction rules, a heterogeneous graph is constructed on the second intermediate data to generate heterogeneous graph data;

[0010] Performing behavioral evolution analysis on the heterogeneous graph data in the structural dimension and the time dimension through the attention mechanism, and dynamically fusing the analysis results using the cross-modal attention mechanism to obtain third intermediate data;

[0011] The third intermediate data is processed by the recommendation model to sort the multiple objects to be identified, and the sorting results are screened according to a preset screening strategy to determine the identification result.

[0012] In some exemplary embodiments, performing entity recognition and topic extraction on the initial text data using a pre-trained large language model includes:

[0013] Performing error correction and cleaning on the initial text data through the generative named entity recognition mechanism of the large language model;

[0014] Entity recognition is performed on the error correction and cleaning results, and topic extraction is performed on the error correction and cleaning results using a topic generation pre-training framework.

[0015] In some exemplary embodiments, the normalizing the first intermediate data using the sentence embedding representation model includes:

[0016] Converting the first intermediate data into a semantic vector of a set dimension through the sentence embedding representation model;

[0017] The semantic vectors are clustered and analyzed using a density clustering algorithm, and different clusters are distinguished according to the clustering results, so as to perform the normalization process.

[0018] In some exemplary embodiments, generating heterogeneous graph data includes:

[0019] During the generation process of the heterogeneous graph data, a graph database is used to manage the generation process of the second intermediate data and the heterogeneous graph data.

[0020] In some exemplary embodiments, performing behavioral evolution analysis on the heterogeneous graph data in the structural dimension and the temporal dimension using an attention mechanism includes:

[0021] Dividing the heterogeneous graph data into at least one time snapshot according to a set rule;

[0022] Using the attention mechanism, perform multi-head weighted aggregation on at least one time snapshot in the structural dimension to obtain at least one local feature result;

[0023] According to the moment corresponding to the least one local feature result, the attention mechanism is used to perform embedding sequence modeling at different moments in the time dimension, the position encoding technology and the causal masking technology are used to ensure the time sequence in the modeling process, and the evolution trend of the multiple objects to be identified over time is determined according to the modeling results.

[0024] In some exemplary embodiments, dynamically fusing analysis results using a cross-modal attention mechanism includes:

[0025] Inputting the at least one time snapshot into the large language model to generate a unified node description;

[0026] Processing the unified node description using a pre-trained model to obtain a semantic embedding;

[0027] The cross-modal attention mechanism is used to dynamically fuse the analysis results with the semantic embedding.

[0028] In some exemplary embodiments, processing the third intermediate data using a recommendation model includes:

[0029] Performing dimensionality reduction and clustering processing on the third intermediate data, thereby preliminarily dividing the plurality of objects to be identified;

[0030] The preliminary division results are input into a pre-trained Bayesian personalized ranking model to calculate scores for the multiple objects to be identified, and the objects are ranked according to the calculation results.

[0031] Based on the same concept, the present application also provides a target object recognition device, comprising:

[0032] The first module is used to obtain initial text data; wherein the initial text data includes text data published by multiple objects to be identified;

[0033] The second module is configured to perform entity recognition and topic extraction on the initial text data using a pre-trained large language model to obtain first intermediate data;

[0034] A third module is configured to input the first intermediate data into a pre-trained sentence embedding representation model, and perform normalization processing on the first intermediate data using the sentence embedding representation model to obtain second intermediate data;

[0035] A fourth module is configured to construct a heterogeneous graph on the second intermediate data according to a set graph network construction rule to generate heterogeneous graph data;

[0036] The fifth module is used to analyze the behavioral evolution of the heterogeneous graph data in the structural dimension and the temporal dimension through the attention mechanism, and dynamically fuse the analysis results using the cross-modal attention mechanism to obtain third intermediate data;

[0037] The sixth module is used to process the third intermediate data through the recommendation model to sort the multiple objects to be identified, filter the sorting results according to a preset filtering strategy, and determine the identification result.

[0038] Based on the same concept, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above methods when executing the computer program.

[0039] Based on the same concept, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement any of the methods described above.

[0040] As can be seen from the above, the present application provides a method, device, electronic device and storage medium for target object identification. The present application realizes named entity recognition of the initial text data through a large language model, and performs topic refinement and extraction through the corresponding model architecture; then, the sentence embedding representation model is used to normalize the synonymous entities, which can effectively reduce data redundancy. Subsequently, the association relationship between users, posts and normalized keywords can be established according to the pre-set, and this can be constructed into heterogeneous graph data, and then the attention mechanism is used to fuse the deep semantics of the text at the structural level and the temporal level to generate a unified node representation; finally, the risk assessment and priority sorting of the identified objects are performed based on the sorting recommendation model, so that potential high-risk target objects can be accurately discovered without relying on the direct social relationship of users. The recognition results can then be used to feed back the method model, so that continuous dynamic updates can be achieved, ensuring that potential high-risk target persons can be quickly and accurately identified in a large-scale dynamic social network environment, providing strong technical support for public safety and platform governance. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] Figure 1 A schematic diagram of the overall architecture of the exemplary method provided in the embodiments of the present application.

[0043] Figure 2 A flowchart of an exemplary method provided in an embodiment of the present application.

[0044] Figure 3 A schematic diagram of the specific process of dynamic fusion and sorting screening according to the exemplary method provided in the embodiments of the present application.

[0045] Figure 4 A schematic diagram of the structure of an exemplary device provided in an embodiment of the present application.

[0046] Figure 5 A schematic diagram of the electronic device structure provided in an embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of this specification more clear, this specification is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0048] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements, objects or method steps that appear before the word cover the elements, objects or method steps listed after the word and their equivalents, without excluding other elements, objects or method steps. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0049] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0050] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0051] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0052] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0053] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0054] As described in the background technology section, currently, in related technologies, research is mainly carried out from the perspective of social network analysis.

[0055] In some embodiments, social network characteristics analysis can be performed on certain network organizations. This embodiment calculates the influence distribution of group members, identifies core users, and analyzes their anonymity characteristics. Simultaneously, it studies the clustering and resource distribution within the community, as well as the relationship between topics and changes over time, providing a theoretical reference for tracking criminal activity.

[0056] In other embodiments, a registration information-based detection model is proposed for malicious accounts on the WeChat platform. This model analyzes account registration features, such as IP addresses and mobile phone numbers, to construct a similar account connectivity graph and mine connected components of suspected malicious accounts, enabling rapid detection of malicious accounts.

[0057] Related technologies also include methods that analyze behavioral patterns. For example, some embodiments propose an unknown attack detection model based on intranet behavior analysis. This embodiment collects intranet information resources, analyzes the behavioral anomaly risk factors of information nodes, and constructs a detection directed graph model to effectively detect unknown attack methods. Other embodiments propose a full-feature information balance modeling method for detecting insider threats. This method analyzes abnormal or inconsistent behavior in log data and uses machine learning algorithms to build a model to effectively detect insider threats.

[0058] It can be seen that in the field of target object discovery, although various methods have achieved certain results, most of the related technologies rely on static analysis models of social relationship networks and cannot effectively reflect the dynamic characteristics of user interactions, posts, and interest evolution over time. This static model often ignores the temporal changes in user behavior, resulting in a significant lag in real-time monitoring and rapid response, which in turn affects the timely identification of target persons. At the same time, related technologies often rely solely on social relationships or single text features to mine target persons, lacking the effective integration of graph structure information and the deep semantics of the text. A single method cannot fully utilize the semantic information hidden in the text, nor can it capture the implicit risks in user behavior through simple keyword matching, resulting in frequent misjudgments or omissions during the identification process.

[0059] In combination with the above actual situation, the embodiment of the present application provides a method for identifying a target object. The present application realizes named entity recognition of the initial text data through a large language model, and performs topic refinement and extraction through the corresponding framework of the model; then, the sentence embedding representation model is used to normalize the synonymous entities, which can effectively reduce data redundancy. Subsequently, the association relationship between users, posts and normalized keywords can be established according to the pre-set, and this can be constructed into heterogeneous graph data, and then the attention mechanism is used to fuse the deep semantics of the text at the structural level and the time level to generate a unified node representation; finally, the risk assessment and priority sorting of the identified objects are performed based on the sorting recommendation model, so that potential high-risk target objects can be accurately discovered without relying on the direct social relationship of users. The recognition results can then be used to feed back the method model, so that continuous dynamic updates can be achieved, ensuring that potential high-risk target persons can be quickly and accurately identified in a large-scale dynamic social network environment, providing strong technical support for public safety and platform governance.

[0060] like Figure 1 As shown, in a specific application scenario, the overall framework of this embodiment can be divided into two parts: "dynamic heterogeneous graph construction" and "target discovery": in the "dynamic heterogeneous graph construction" part, the original data (initial text data) can first be subjected to noise cleaning, entity recognition, topic extraction, and entity normalization to construct a compact and highly connected heterogeneous graph; in the "target discovery" part, target person mining can be performed based on the heterogeneous graph, and a ranked list of suspicious target objects can be output through a comprehensive analysis of node features, relationship structure, and spatiotemporal evolution information. The results can then be returned to the data source for continuous dynamic updating.

[0061] After that, for the specific workflow, Figure 2 A flow chart of an exemplary method provided in an embodiment of the present application is shown.

[0062] like Figure 2 As shown, the target object recognition method exemplarily proposed in the embodiment of the present application includes the following steps.

[0063] Step 202: Acquire initial text data; wherein the initial text data includes text data published by multiple objects to be identified.

[0064] In this step, the initial text data is used for target object recognition. The text here can be provided by the user through the corresponding port when using the corresponding recognition terminal, or it can be imported in batches through a table or link. This initial text data generally contains multiple objects to be recognized, such as articles published by different objects to be recognized (such as users), news written, publicly available forum posts, dynamics, etc., and some initial text data may also contain text content exchanged between multiple objects to be recognized.

[0065] Afterwards, the core purpose of this embodiment is to identify the target object among these objects to be identified. Here, the "target object" is defined as a high-risk user group that may engage in illegal activities or have a significant adverse impact on platform public opinion and social security. It explores how to use massive social media data and deeply mine user behavior to achieve accurate identification and targeted monitoring of potential threat users, thereby providing strong support for public safety and platform governance.

[0066] Step 204 : Use a pre-trained large language model to perform entity recognition and topic extraction on the initial text data to obtain first intermediate data.

[0067] In this step, first, a large language model (LLM) is an artificial intelligence model based on deep learning that is designed to understand and generate human language. It usually contains billions or even hundreds of billions of parameters and is trained with massive amounts of text data. The core capabilities of these models include natural language processing (NLP) tasks such as text generation, machine translation, question answering, sentiment analysis, etc. Here, this step can be performed using a generative large language model (GPT-type model) or a deep pre-trained model (BERT variant) in a large language model, without specific limitation.

[0068] In this embodiment, after receiving the initial text data from an external platform or user input, it is first necessary to extract effective information from the initial text data. Here, it is mainly necessary to accurately identify the entities therein, and it is also necessary to identify and extract the topic themes of these entities.

[0069] In specific application scenarios, if the initial text data itself is relatively regular and conforms to the corresponding specifications, it can be directly input into the large language model for subsequent recognition and extraction. However, in more common scenarios, due to the widespread noise and multilingual mixing problems in the initial text data from external platforms, specifically, the text usually initially contains abbreviations, irregular spellings, emoticons, and mixed languages. Traditional entity recognition methods based on dictionaries or fixed rules are often difficult to cope with. Traditional text processing methods have low recognition accuracy when dealing with typos, abbreviations, and colloquial expressions that are prevalent in social media, and can easily cause the same entity to be recognized multiple times, resulting in a large number of redundant nodes. This not only destroys the structural compactness of the data, but also reduces the effectiveness of the graph structure in subsequent analysis.

[0070] In this way, in some embodiments, the generative named entity recognition (GPTNER) mechanism of the large language model can be used as the core to perform adaptive error correction on the input noisy text. Specifically, the generative named entity recognition mechanism will automatically identify typos, abbreviations and colloquial expressions in the text, and perform error correction and cleaning based on the context information to ensure that key information such as names of people, places, organizations and events can be accurately extracted, so that the large language model can be used to more conveniently perform entity recognition. Afterwards, for the error correction and cleaning results, the topic generative pre-training framework (TopicGPT) technology can be further referenced to extract topics from the text content and generate semantically condensed and compact topic phrases. This step not only overcomes the problem that traditional statistical methods such as LDA are prone to topic generalization when processing short texts, but can also dynamically adapt to the expression characteristics of texts in different fields, laying a solid foundation for subsequent structured representation. That is, in some embodiments, the use of a pre-trained large language model to perform entity recognition and topic extraction on the initial text data includes: performing error correction and cleaning on the initial text data through the generative named entity recognition mechanism of the large language model; performing entity recognition on the error correction and cleaning results, and using a topic generation pre-training framework to perform topic extraction on the error correction and cleaning results.

[0071] In more specific application scenarios, a generative large language model (GPT-like model) or a deep pre-trained model (BERT variant) can be introduced with prompt guidance. After the post content of the target is segmented and preliminarily cleaned, the model automatically outputs entities such as person names (PERSON), place names (LOCATION), organizations (ORG), and events (EVENT) that may appear in the post text, and corrects noisy spellings. For example, if a post mixed with English and pinyin reads: "On Pacifc regin, USN zmlt ddg1000 conducts military exercises," traditional methods may not be able to correctly identify the place name "Pacific region" or the ship entity "Zumwalt (DDG-1000)" due to language switching and spelling errors. In this embodiment, by adding instructions to the large language model, such as "Please identify keywords such as ships, troops and place names from the following text, and correct obvious spelling errors", the model can generate annotations such as "Pacificregion@@LOCATION##" and "Zumwalt(DDG-1000)@@WEAPON##", thereby extracting accurate entities from informal and easily confused texts. At the same time, in order to assist in higher-level semantic analysis, the present invention also designs a topic extraction function, which will first cluster the post text into short texts, and then let the large language model generate topic names based on several keywords, and merge repeated or similar topics, aggregating "multi-domain military exercises" and "cross-service joint exercises" into "multi-domain collaborative exercises". The entity annotations and topic tags produced by this module will provide basic semantic information for subsequent modules, thereby achieving preliminary structuring and simplification of massive high-noise social texts, facilitating the retention of sufficient key information in subsequent graph modeling and spatiotemporal analysis.

[0072] Finally, the entity recognition results and topic extraction results are output to form the first intermediate data. In a specific application scenario, the entity recognition results can be output in the form of an entity list, and the topic extraction results can be output in the form of topic tags.

[0073] Step 206: Input the first intermediate data into a pre-trained sentence embedding representation model, and use the sentence embedding representation model to normalize the first intermediate data to obtain second intermediate data.

[0074] In this step, since the initial text data in some embodiments, such as the initial text data provided by social media platforms, often contains the same entity in different forms, if normalization is not performed, there will be a large number of redundant nodes in the graph structure subsequently constructed, thereby reducing data quality and network connectivity. Afterwards, the sentence embedding representation model can be a SentenceBERT model, which is a pre-trained model based on BERT and is designed for generating sentence embeddings. It is mainly used for tasks such as calculating semantic similarity, text clustering, and information retrieval. It optimizes the generation of sentence embeddings to make the calculation of the similarity between sentence pairs more efficient.

[0075] In some embodiments, a pre-trained sentence embedding representation model can be used to convert entities and topics extracted from the initial text data (i.e., the first intermediate data) into fixed-dimensional semantic vectors, and then density clustering algorithms such as the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm and the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm can be used to perform cluster analysis on these vectors. Through clustering, different expressions with similar semantics can be automatically merged into a unified identifier. This process not only reduces redundant data and improves the consistency of information, but also ensures that text data from different sources and formats can be uniformly expressed in the same semantic space, providing refined input for the subsequent construction of the graph network. That is, in some embodiments, the specific process of normalization processing, wherein the normalization processing of the first intermediate data using the sentence embedding representation model includes: converting the first intermediate data into semantic vectors of a set dimension using the sentence embedding representation model; performing cluster analysis on the semantic vectors using a density clustering algorithm, and distinguishing different clusters based on the clustering results, thereby performing the normalization processing. The set dimension is the aforementioned fixed dimension, for example, it can be 768 dimensions.

[0076] In a more specific application scenario, after obtaining a preliminary list of entities and topic labels, in order to prevent entities that actually point to the same object from appearing repeatedly due to spelling variations, capitalization differences, special abbreviations, etc., automatic merging and deduplication can be performed based on the aforementioned first intermediate data. Here, two major technologies, text vectorization and unsupervised clustering, can be combined: First, a sentence-level or phrase-level vector representation can be generated for each entity name, and a sentence embedding representation model, a pre-trained model that can process short texts, can be used to map the forms "zmlt ddg-1000" and "Zumwalt (DDG-1000)" into a high-dimensional semantic space; second, based on the cosine similarity values ​​between entities, a density clustering algorithm such as DBSCAN or HDBSCAN is executed to automatically cluster entities with similarities exceeding a certain threshold into the same "synonymous cluster", and then a "representative entity name" is selected from the cluster through a weighted strategy as the normalized unified name. For example, when analyzing posts related to a certain country's navy, if there are five or six ways of writing such as "ZumwltDDG1000" and "Zmlt(DDG-1000)", but the similarity of these writings in the semantic vector space is greater than 0.85, this step can be used to merge them into a standard entity node called "Zumwalt(DDG-1000)", reducing redundancy or confusion in the subsequent graph construction process. In addition to abbreviations, this step can also solve the problem of mixed pinyin in some writings. For example, "BeiJing" and "Beijing" can be attributed to the same place in a certain context. Finally, through named entity normalization, the impact of "synonymous multiple writing" on graph nodes can be effectively reduced, and it is ensured that when entities with the same meaning appear in large-scale texts, they can point to a unified node identifier, which significantly improves the compactness and connectivity of the heterogeneous graph structure in the subsequent steps.

[0077] Finally, the normalized clustering results are output to form the second intermediate data.

[0078] Step 208: construct a heterogeneous graph on the second intermediate data according to the set graph network construction rules to generate heterogeneous graph data.

[0079] In this step, the main purpose is to convert the normalized key information (second intermediate data) into structured data and construct a graph network with multiple types of nodes and multiple relationship edges.

[0080] In the specific implementation process, users (i.e., objects to be identified), posts, and keywords (including named entities and topics) extracted from the second intermediate data can first be treated as different types of nodes. Then, based on the publishing relationships between users and posts, the mention relationships between posts and keywords, and the attribute associations between users and keywords, corresponding edges are established to form graph network construction rules. The heterogeneous graph constructed using these graph network construction rules not only reflects direct social interactions between users but also captures the implicit semantic connections in the text content. The resulting heterogeneous graph is heterogeneous graph data.

[0081] In some embodiments, to ensure efficient data storage and query, the entire graph network can be managed using a graph database (e.g., a Neo4j database, etc.), so that the nodes and edges in the graph network have high query efficiency and visualization capabilities, providing a solid data foundation for subsequent spatiotemporal feature fusion. That is, in some embodiments, generating heterogeneous graph data includes: during the generation of the heterogeneous graph data, using a graph database to manage the second intermediate data and the generation process of the heterogeneous graph data.

[0082] In more specific application scenarios, after completing entity recognition, topic extraction and normalization, users, posts, keywords (including both previously generated entities and abstract topics) can be stored in a structured heterogeneous graph. In this step, three main types of nodes can be mapped to the corresponding graph database. Taking the Neo4j database as an example, three tags, User, Post, and Keyword, can be created in Neo4j: the User node stores the user's (the object to be identified) ID, name, place of residence, etc.; the Post node records the post text, publishing timestamp, likes and comments data; the Keyword node represents different types of entities or topics, such as "Zumwalt (DDG-1000)", "cross-service exercise", etc., and is classified according to tags (Weapon, Event, Topic, etc.). Edges can then be inserted based on the type of relationship. For example, a "POSTED" relationship can be established between a user and a post, a "MENTIONS" relationship can be established between a post and a keyword, a "FRIEND" relationship can be established between users who are friends or relatives, and an "ASSOCIATED_WITH" relationship can be created if a user mentions a keyword in their profile. For example, if user U1's profile states "working at XX Institute," user U1 will be connected to "XX Institute@@ORG##" in the heterogeneous graph. If a new user U2 subsequently claims to work at XX Institute, U2 will also be connected to the same keyword node, forming a potential colleague relationship, which facilitates the subsequent mining of higher-level semantics between the two. This heterogeneous network, composed of multiple node types and multiple edge types, is particularly suitable for the complex connections in social scenarios, avoiding the loss of type distinctions caused by cramming all information into a homogeneous graph. Later, in some embodiments, time information can be used when constructing a heterogeneous graph to mark posts, user attribute changes, etc. with their respective timestamps for use in downstream spatiotemporal analysis and relationship evolution tracking.

[0083] In step 210 , the attention mechanism is used to perform behavioral evolution analysis on the heterogeneous graph data in the structural dimension and the time dimension, and the cross-modal attention mechanism is used to dynamically fuse the analysis results to obtain third intermediate data.

[0084] In this step, in order to better capture the changing trends of the behavior of the object to be identified and the text content in the time dimension, this step can be used to combine graph structure attention and time series information. At the same time, it can also be combined with the text embedding of previous keywords to form a complete meaning representation.

[0085] It can be seen that this step aims to integrate graph structure information, time evolution features and deep text semantics to fully reflect the dynamic changes in the behavior and interests of the object to be identified. Afterwards, the attention mechanism is a neural network architecture that simulates human attention, which is designed to help the model process input information more effectively, especially long sequence data. It enables the model to focus on the most relevant parts of the current task by dynamically assigning weights, thereby improving the performance and efficiency of the model. In specific application scenarios, such as Figure 3 As shown, heterogeneous graph data can first be partitioned according to set time intervals to form multiple time snapshots. Within each time snapshot, a structural attention layer can be used to perform multi-head weighted aggregation on information about each node and its neighborhood in the graph, extracting local features of the current graph structure. Subsequently, a temporal attention layer can be used to model the embedding sequence of nodes at different time points. Positional encoding and causal masking techniques are used to ensure the correct temporal order, thereby capturing the temporal evolution of the behavior of the target object, completing evolutionary analysis, and forming a structural / temporal embedding. The structural attention layer is a connection layer whose attention mechanism focuses on the structural dimension, while the temporal attention layer is a connection layer whose attention mechanism focuses on the temporal dimension. A cross-modal attention mechanism can then be used to dynamically fuse the multidimensional representation of the graph structure with the node description, automatically adjusting the weights of each modality to generate a unified node representation that combines structural, temporal, and semantic features. This fused representation can reflect both real-time social interactions and the long-term behavioral patterns of the target object, providing comprehensive feature support for accurate identification of the target object. That is, in some embodiments, the behavioral evolution analysis of the heterogeneous graph data in the structural dimension and the time dimension is performed through the attention mechanism, including: dividing the heterogeneous graph data into at least one time snapshot according to the set rules; using the attention mechanism to perform multi-head weighted aggregation on at least one time snapshot in the structural dimension to obtain at least one local feature result; according to the moment corresponding to the at least one local feature result, using the attention mechanism to perform embedded sequence modeling at different moments in the time dimension, using position encoding technology and causal masking technology to ensure the time order in the modeling process, and determining the evolution trend of the multiple objects to be identified over time based on the modeling results.

[0086] In some embodiments, as Figure 3As shown in the figure, in order to further enhance the expression of text information, the aforementioned large language model can be introduced at this stage to generate a unified node description after inputting the time snapshot. This process integrates the personal information, post content and normalized keywords of the object to be identified into a structured text, and then uses pre-trained models such as BERT (Bidirectional Encoder Representations from Transformers) to generate deep semantic embedding. Afterwards, in the cross-modal attention mechanism ( Figure 3 When dynamically fusing the multi-dimensional representation of the graph structure with the node description (corresponding to the multi-head self-attention modal cross-fusion module), the node description is replaced with the deep semantic embedding generated by this embodiment for dynamic fusion. That is, in some embodiments, the dynamic fusion of the analysis results using the cross-modal attention mechanism includes: inputting the at least one time snapshot into the large language model to generate a unified node description; processing the unified node description using a pre-trained model to obtain a semantic embedding; and dynamically fusing the analysis results with the semantic embedding using the cross-modal attention mechanism.

[0087] In a more specific application scenario, on the one hand, the objects and related topics that users (the objects to be identified) follow on social platforms may shift over time. Using only static graph neural networks often makes it difficult to reflect differences between long-term and short-term periods. On the other hand, post texts are inherently semantically rich, and keywords such as "aircraft carrier deployment" and "long-range strike" indicate the user's focus on the military field. This embodiment first aggregates the features of each user node's neighbor nodes using structural attention at the current snapshot. Then, at the temporal level, causal attention is used to fuse the user's historical state with the most recent neighbor interactions to characterize the user's dynamic interests. At the same time, the semantic vector embedding (BERT encoding) of the post keywords is incorporated into the aggregation, allowing the understanding of the node "Zumwalt (DDG-1000)" to be derived not only from its edge measurements in the graph, but also from its textual description of attributes such as "stealth destroyer," thereby enhancing the recognition of technical or organizational characteristics. For example, suppose user U1 posts frequently about "naval base exercises" in January, then focuses on "maritime cruises" and a specific ship in February. These temporal changes in keyword switching and structural connectivity are comprehensively reflected in U1's final embedding representation, allowing subsequent judgments to identify U1's shift in focus to a new topic. This invention, by integrating structural, temporal, and textual information, enables each user to obtain a more realistic and evolving vector representation, providing a reliable dynamic semantic basis for identifying key users.

[0088] Finally, the fusion result obtained by using the cross-modal attention mechanism is the third intermediate data, which can be a fusion representation. In specific application scenarios, the third intermediate data can usually be a unified node representation.

[0089] Step 212: Process the third intermediate data using a recommendation model to sort the multiple objects to be identified, filter the sorting results according to a preset filtering strategy, and determine the identification result.

[0090] In this step, the third intermediate data, after being fused in step 210, can generally be represented as a unified node. Finally, a corresponding ranking recommendation algorithm can be used to mine and rank the risks of all objects to be identified. That is, in specific applications, by combining the user vectors of the objects to be identified, the associations between posts and keywords, and the structural features of the objects to be identified in the network, we can identify and screen out target objects that may be high-risk or require priority attention. Subsequently, the recommendation model corresponds to the ranking recommendation algorithm. The ranking recommendation algorithm is a very critical link in the recommendation system. Its purpose is to perform a secondary scoring and ranking of the results of the recall phase to obtain a more accurate recommendation list.

[0091] In some embodiments, after obtaining the third intermediate data, if the dimension of the third intermediate data meets the requirements of the corresponding recommendation model, the recommendation model can be used to calculate the risk score of each node in the third intermediate data, that is, each object to be identified. The specific calculation rules can be designed according to the specific application scenario. For example, after detecting certain keywords, the corresponding object to be identified is considered to be more dangerous, and its corresponding score will be increased. After calculating the risk scores of multiple objects to be identified, the objects to be identified can be screened according to a preset screening strategy, such as screening according to a preset risk threshold or directly extracting the top N objects to be identified, thereby forming a final recognition result. The recognition result can be presented in text or table form to record the recognition result, which can be a high-risk target object that needs special attention. After forming a ranked list of target objects and other recognition results, the ranked list can be displayed on the current terminal or transmitted, for example, for real-time monitoring and further intervention by departments related to security supervision and risk warning, to ensure that potential high-risk target objects can be quickly and accurately identified in a large-scale dynamic social network environment.

[0092] In some embodiments, since the third intermediate data generally has a high dimensionality, it is necessary to perform corresponding dimensionality reduction processing. At the same time, in order to facilitate the subsequent calculation of the recommendation model, the third intermediate data can be clustered first, and the nodes with higher similarity can be clustered, so that the object groups with similar behavior and interest characteristics can be preliminarily divided. Afterwards, based on considerations such as processing implicit feedback data, optimizing sorting tasks, and improving recommendation quality, a recommendation model such as Bayesian Personalized Ranking (BPR) can be selected in the recommendation model to calculate the risk score of each object to be identified. That is, in some specific application scenarios, the third intermediate data is first subjected to dimensionality reduction and clustering processing to preliminarily divide the object groups with similar behavior and interest characteristics; then, a recommendation model such as Bayesian Personalized Ranking (BPR) is used to calculate the risk score of each object to be identified, and the objects to be identified are output as key target objects based on a preset risk threshold or by directly extracting the top K ranked objects to be identified. That is, in some embodiments, the processing of the third intermediate data by the recommendation model includes: performing dimensionality reduction and clustering processing on the third intermediate data to preliminarily divide the multiple objects to be identified; inputting the preliminary division results into a pre-trained Bayesian personalized ranking model to calculate scores for the multiple objects to be identified, and sorting them according to the calculation results.

[0093] In more specific application scenarios, such as Figure 3As shown, after extracting user representations (third intermediate data, unified node representations) based on the cross-modal attention mechanism, scores can be assigned to each target object based on specific target criteria (e.g., whether the target object frequently appears in extreme topics or highly sensitive areas, or whether its proximity and interaction strength with several known risk objects are significantly increased). This scoring outputs a ranked list, which then transmits the unified node representation to the recommendation module for ranking learning using a recommendation model such as BPR. Ultimately, a recommended list of target objects is formed. Specifically, a decision function can be set for each target object to analyze its neighborhood in the graph to determine factors such as "how many high-risk keywords it involves," "how many similar users follow it," and "the distribution of similar topics in recent posts." If the target object exceeds a predetermined threshold in certain dimensions, its score will be increased accordingly. A ranking algorithm is then used to sort all the targets, and the top-ranked ones are selected as "candidate target objects." For example, if user U2's posts repeatedly contain sensitive entity or organization names, and their relationships with labeled suspect users increase, user U2's ranking in the list provided by the present invention will rise rapidly, prompting security personnel or administrators to conduct a more in-depth review. Furthermore, to enhance the interpretability of the judgment results, the method also lists the links between user U2 and key keywords or nodes, and displays their posting timeline, allowing human reviewers to more intuitively see when and on what topics the apparent high-risk behavior began. This comprehensive judgment and ranking process allows for efficient screening of the most noteworthy targets within a vast, heterogeneous network. Subsequent versions will also allow for continuous access to their new posts or friend data, strengthening dynamic monitoring capabilities and providing precise and traceable technical means for public safety and platform security governance.

[0094] Finally, the recognition results can be output, for example, they can be displayed on a corresponding device to provide the operator with corresponding feedback. Of course, in other embodiments, the output method of the recognition results may not be limited to output display, and it may also be used to store, display, use or reprocess the recognition results. The specific output method of the recognition results can be flexibly selected according to different application scenarios and implementation needs.

[0095] Specifically, for example, in an application scenario where the method of this embodiment is executed on a single device, the recognition result can be directly output in a displayed manner on the display component (display, projector, etc.) of the current device, so that the operator of the current device can directly see the content of the recognition result from the display component.

[0096] For another example, in an application scenario where the method of this embodiment is executed on a system composed of multiple devices, the recognition results can be sent to other preset devices serving as recipients within the system, i.e., synchronization terminals, through any data communication method (wired connection, NFC, Bluetooth, Wi-Fi, cellular mobile network, etc.), so that the synchronization terminals can perform subsequent processing on them. Optionally, the synchronization terminal can be a preset server, which is generally located in the cloud and serves as a data processing and storage center, capable of storing and distributing the recognition results; wherein the recipients of the distribution are terminal devices, and the holders or operators of these terminal devices can be managers, supervisors, maintenance personnel, personnel from relevant departments, etc. of the identification system.

[0097] For another example, in an application scenario where the method of this embodiment is executed on a system composed of multiple devices, the recognition result can be sent directly to a preset terminal device through any data communication method. The terminal device can be one or more of the devices listed in the preceding paragraphs.

[0098] It can be seen from the above embodiments that the embodiments of the present application provide a method for identifying a target object. The present application realizes named entity recognition of the initial text data through a large language model, and performs topic refinement and extraction through the corresponding framework of the model; then, the sentence embedding representation model is used to normalize the synonymous entities, which can effectively reduce data redundancy. Subsequently, the association relationship between users, posts and normalized keywords can be established according to the pre-set, and this can be constructed into heterogeneous graph data, and then the attention mechanism is used to fuse the deep semantics of the text at the structural level and the temporal level to generate a unified node representation; finally, the risk assessment and priority sorting of the identified objects are performed based on the sorting recommendation model, so that potential high-risk target objects can be accurately discovered without relying on the direct social relationship of users. The recognition results can then be used to feed back the method model, so that continuous dynamic updates can be achieved, ensuring that potential high-risk target persons can be quickly and accurately identified in a large-scale dynamic social network environment, providing strong technical support for public safety and platform governance.

[0099] From the above, it can be seen that the purpose of this application is to provide a method for mining target persons in heterogeneous graphs with spatiotemporal semantic fusion. This method can use advanced large language models to perform adaptive error correction, named entity recognition and topic extraction on noisy texts in social media, and use semantic clustering technology to realize the normalization of synonymous or deformed entities, thereby constructing a compact and highly connected heterogeneous graph; at the same time, the dynamic evolution of user behavior and interests is captured through the structural attention layer and the temporal attention layer, and combined with the deep semantic embedding of the text, the effective fusion of graph structure information and text information is realized, and finally the purpose of accurately identifying and real-time screening of high-risk target objects is achieved, providing strong technical support for public safety and platform governance.

[0100] In specific application scenarios, this application can quickly extract the core semantics of the object to be identified in a high-noise, redundant text environment and construct multi-type entity nodes. By using adaptive error correction and noise filtering, the method can eliminate low-value content such as advertisements, repeated topics and typos, and accurately extract named entities and topic information in posts through a large language model. Subsequently, synonymous and near-synonymous expressions are normalized by combining technologies such as Sentence-BERT and density clustering to avoid the same object being modeled repeatedly multiple times; this not only reduces the time cost of text processing, but also ensures that the key texts ultimately retained by users are highly representative, laying a good data foundation for subsequent heterogeneous graph modeling and spatiotemporal semantic fusion. Afterwards, this application can also efficiently discover high-risk target objects and complete priority sorting without relying on a single social relationship. At the same time, a heterogeneous graph model can be used to integrate users, posts and keywords into a unified dynamic network, and identify the temporal changes in the behavior of the object to be identified through structural attention and temporal attention, and then superimpose deep semantic information to generate a unified node representation. Based on this fused representation, the method utilizes a ranked recommendation algorithm (BPR) to rapidly calculate risk scores in a large-scale user and post environment, while also leveraging multidimensional association paths to enhance interpretability. Ultimately, field tests demonstrate that this embodiment can accurately distinguish between general content and sensitive topics that potentially threaten social security, allowing rapid identification of targets worthy of focused monitoring in a dynamic environment.

[0101] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of the embodiment of the present application can also be applied in a distributed scenario and completed by multiple devices working together. In the case of such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method described.

[0102] It should be noted that the above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0103] Based on the same concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a target object recognition device.

[0104] refer to Figure 4 , the target object recognition device includes:

[0105] The first module 410 is configured to obtain initial text data, wherein the initial text data includes text data published by a plurality of objects to be identified.

[0106] The second module 420 is used to use a pre-trained large language model to perform entity recognition and topic extraction on the initial text data to obtain first intermediate data.

[0107] The third module 430 is used to input the first intermediate data into a pre-trained sentence embedding representation model, and use the sentence embedding representation model to normalize the first intermediate data to obtain second intermediate data.

[0108] The fourth module 440 is configured to construct a heterogeneous graph on the second intermediate data according to a set graph network construction rule to generate heterogeneous graph data.

[0109] The fifth module 450 is used to perform behavioral evolution analysis on the heterogeneous graph data in the structural dimension and the time dimension through the attention mechanism, and dynamically fuse the analysis results using the cross-modal attention mechanism to obtain third intermediate data.

[0110] The sixth module 460 is configured to process the third intermediate data using a recommendation model to sort the multiple objects to be identified, filter the sorting results according to a preset filtering strategy, and determine the identification result.

[0111] In some exemplary embodiments, the second module 420 is further configured to:

[0112] Performing error correction and cleaning on the initial text data through the generative named entity recognition mechanism of the large language model;

[0113] Entity recognition is performed on the error correction and cleaning results, and topic extraction is performed on the error correction and cleaning results using a topic generation pre-training framework.

[0114] In some exemplary embodiments, the third module 430 is further configured to:

[0115] Converting the first intermediate data into a semantic vector of a set dimension through the sentence embedding representation model;

[0116] The semantic vectors are clustered and analyzed using a density clustering algorithm, and different clusters are distinguished according to the clustering results, so as to perform the normalization process.

[0117] In some exemplary embodiments, the fourth module 440 is further configured to:

[0118] During the generation process of the heterogeneous graph data, a graph database is used to manage the generation process of the second intermediate data and the heterogeneous graph data.

[0119] In some exemplary embodiments, the fifth module 450 is further configured to:

[0120] Dividing the heterogeneous graph data into at least one time snapshot according to a set rule;

[0121] Using the attention mechanism, perform multi-head weighted aggregation on at least one time snapshot in the structural dimension to obtain at least one local feature result;

[0122] According to the moment corresponding to the least one local feature result, the attention mechanism is used to perform embedding sequence modeling at different moments in the time dimension, the position encoding technology and the causal masking technology are used to ensure the time sequence in the modeling process, and the evolution trend of the multiple objects to be identified over time is determined according to the modeling results.

[0123] In some exemplary embodiments, the fifth module 450 is further configured to:

[0124] Inputting the at least one time snapshot into the large language model to generate a unified node description;

[0125] Processing the unified node description using a pre-trained model to obtain a semantic embedding;

[0126] The cross-modal attention mechanism is used to dynamically fuse the analysis results with the semantic embedding.

[0127] In some exemplary embodiments, the sixth module 460 is further configured to:

[0128] Performing dimensionality reduction and clustering processing on the third intermediate data, thereby preliminarily dividing the plurality of objects to be identified;

[0129] The preliminary division results are input into a pre-trained Bayesian personalized ranking model to calculate scores for the multiple objects to be identified, and the objects are ranked according to the calculation results.

[0130] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0131] The apparatus of the above embodiment is used to implement the corresponding target object recognition method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0132] Based on the same concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the computer program, the target object recognition method as described in any of the above embodiments is implemented.

[0133] Figure 5 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.

[0134] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0135] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0136] The input / output interface 1030 is used to connect an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0137] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).

[0138] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).

[0139] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0140] The electronic device of the above embodiment is used to implement the corresponding target object recognition method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0141] Based on the same concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the target object recognition method described in any of the above embodiments.

[0142] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0143] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the target object recognition method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0144] Based on the same concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides a computer program product comprising computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the target object recognition method. Corresponding to the execution subject corresponding to each step in each embodiment of the target object recognition method, the processor executing the corresponding step can belong to the corresponding execution subject.

[0145] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the target object recognition method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0146] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0147] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.

[0148] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.

[0149] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.

Claims

1. A target object recognition method, characterized in that: include: Acquire initial text data; wherein the initial text data includes text data published by multiple objects to be identified; Using a pre-trained large language model to perform entity recognition and topic extraction on the initial text data to obtain first intermediate data; Inputting the first intermediate data into a pre-trained sentence embedding representation model, and normalizing the first intermediate data using the sentence embedding representation model to obtain second intermediate data; According to the set graph network construction rules, a heterogeneous graph is constructed on the second intermediate data to generate heterogeneous graph data; Performing behavioral evolution analysis on the heterogeneous graph data in the structural dimension and the time dimension through the attention mechanism, and dynamically fusing the analysis results using the cross-modal attention mechanism to obtain third intermediate data; The third intermediate data is processed by the recommendation model to sort the multiple objects to be identified, and the sorting results are screened according to a preset screening strategy to determine the identification result.

2. The method according to claim 1, characterized in that The method of using a pre-trained large language model to perform entity recognition and topic extraction on the initial text data includes: Performing error correction and cleaning on the initial text data through the generative named entity recognition mechanism of the large language model; Entity recognition is performed on the error correction and cleaning results, and topic extraction is performed on the error correction and cleaning results using a topic generation pre-training framework.

3. The method according to claim 1, characterized in that The normalizing the first intermediate data by using the sentence embedding representation model includes: Converting the first intermediate data into a semantic vector of a set dimension through the sentence embedding representation model; The semantic vectors are clustered and analyzed using a density clustering algorithm, and different clusters are distinguished according to the clustering results, so as to perform the normalization process.

4. The method according to claim 1, wherein Generating heterogeneous graph data includes: During the generation process of the heterogeneous graph data, a graph database is used to manage the generation process of the second intermediate data and the heterogeneous graph data.

5. The method according to claim 1, wherein The behavioral evolution analysis of the heterogeneous graph data in the structural dimension and the time dimension by using the attention mechanism includes: Dividing the heterogeneous graph data into at least one time snapshot according to a set rule; Using the attention mechanism, perform multi-head weighted aggregation on at least one time snapshot in the structural dimension to obtain at least one local feature result; According to the moment corresponding to the least one local feature result, the attention mechanism is used to perform embedding sequence modeling at different moments in the time dimension, the position encoding technology and the causal masking technology are used to ensure the time sequence in the modeling process, and the evolution trend of the multiple objects to be identified over time is determined according to the modeling results.

6. The method according to claim 5, characterized in that The cross-modal attention mechanism is used to dynamically fuse the analysis results, including: Inputting the at least one time snapshot into the large language model to generate a unified node description; Processing the unified node description using a pre-trained model to obtain a semantic embedding; The cross-modal attention mechanism is used to dynamically fuse the analysis results with the semantic embedding.

7. The method according to claim 1, characterized in that The processing of the third intermediate data by using the recommendation model includes: Performing dimensionality reduction and clustering processing on the third intermediate data, thereby preliminarily dividing the plurality of objects to be identified; The preliminary division results are input into a pre-trained Bayesian personalized ranking model to calculate scores for the multiple objects to be identified, and the objects are ranked according to the calculation results.

8. A target object recognition device, characterized in that: include: The first module is used to obtain initial text data; wherein the initial text data includes text data published by multiple objects to be identified; The second module is configured to perform entity recognition and topic extraction on the initial text data using a pre-trained large language model to obtain first intermediate data; A third module is configured to input the first intermediate data into a pre-trained sentence embedding representation model, and perform normalization processing on the first intermediate data using the sentence embedding representation model to obtain second intermediate data; A fourth module is configured to construct a heterogeneous graph on the second intermediate data according to a set graph network construction rule to generate heterogeneous graph data; The fifth module is used to analyze the behavioral evolution of the heterogeneous graph data in the structural dimension and the temporal dimension through the attention mechanism, and dynamically fuse the analysis results using the cross-modal attention mechanism to obtain third intermediate data; The sixth module is used to process the third intermediate data through the recommendation model to sort the multiple objects to be identified, filter the sorting results according to a preset filtering strategy, and determine the identification result.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to implement the method according to any one of claims 1 to 7.