Intelligent natural language processing and text depth analysis method
By constructing a data layout and target disk, and utilizing identity tags and counters for moving points and marker points, the problem of inflexible target analysis in existing technologies is solved, achieving high efficiency and accuracy in text analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN JINGSHENG HUIZHI TECHNOLOGY CO LTD
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-24
AI Technical Summary
Existing language processing and deep text analysis methods struggle to quickly extract relevant data from text data based on different analytical task objectives, impacting the accuracy and efficiency of the analysis.
The system constructs a data layout and target disk, extracts data clusters from the data layout by establishing relationships, and inputs them into a trained analysis model for analysis. It utilizes the identity tags and counters of moving points and marker points to manage data, enabling flexible data selection and efficient analysis.
It improves the accuracy and efficiency of text analysis, adapts to different analysis objectives, reduces redundant analysis, and improves the efficiency of data block analysis.
Smart Images

Figure CN121920350A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of language processing technology, specifically to an intelligent natural language processing and deep text analysis method. Background Technology
[0002] With the rapid development of information technology, natural language processing technology has made significant progress in the field of text analysis. However, existing language processing and deep text analysis methods are difficult to use quickly to extract relevant data from text data according to different analysis task objectives, thus affecting the accuracy and efficiency of text analysis.
[0003] To address these issues, we propose an intelligent natural language processing and deep text analysis method. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent natural language processing and deep text analysis method to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an intelligent natural language processing and deep text analysis method, the method comprising the following steps: Obtain the text data to be analyzed, and construct a data layout corresponding to the text data to be analyzed; The analysis task objective is to acquire text data. Based on the analysis task objective, a target disk is constructed. The target disk consists of multiple moving points, each corresponding to a different analysis task objective. Establish the relationship between the target data panel and the data layout, and extract the data clusters corresponding to the target points from the data layout based on the relationship; The data clusters are input into the trained analysis model, and the data clusters are analyzed based on the analysis model to obtain the text analysis results corresponding to the analysis task objectives.
[0006] Preferably, the steps of acquiring the text data to be analyzed and constructing a data layout corresponding to the text data to be analyzed include: Obtain the text data to be analyzed, divide the text data into multiple data blocks, and set station points for the corresponding multiple data blocks; Communication channels are established sequentially between multiple outposts, and multiple data areas are set for each outpost. Each data area consists of multiple data points arranged sequentially. Multiple sub-data blocks within a data block are arranged sequentially on data points, and these data points are then sorted according to the order in which the data blocks are divided to obtain the data layout.
[0007] Preferably, the step of constructing a target disk based on the analysis task objective of acquiring text data includes: Get multiple analysis task targets corresponding to the same text data source, set a corresponding moving point for each analysis task target, and store multiple moving points in each moving point; Within each mobile point, multiple marker points and multiple docking points are configured. For each marker point, multiple applicable tags are configured. The applicable tags are the identity tags of other mobile points that can be enabled by the content marked by the marker point. Establish a communication connection between the marker point and the docking point, and set the trigger condition, wherein the trigger condition is that the applicable tag corresponding to the marker point contains the identity tag of the mobile point corresponding to the docking point; The target disk is obtained by connecting the movement points corresponding to multiple analysis task objectives.
[0008] Preferably, the step of configuring multiple marker points within each moving point and configuring multiple applicable tags for each marker point includes: Configure a unique identity tag for each mobile point where the analysis task target is located, and store multiple identity tags corresponding to different analysis task targets at each mobile point to obtain an identity tag library; The current location of the mobile point and its corresponding relationship are obtained as the tagging information. Based on the tagging information, multiple identity tags are selected from the identity tag library, and the selected identity tags are temporarily assigned as the applicable tags for the current tagging point. A counter is set for each marker point, and the initial value of the counter is determined based on the number of applicable tag types configured for the marker point.
[0009] Preferably, the step of establishing the association between the target disk and the data layout, and extracting the data clusters corresponding to the target points from the data layout based on the association includes: Obtain the mobile points on the target disk, and deploy multiple identical mobile points simultaneously on a single mobile point; divide the data layout composed of target text data into regions to form multiple analysis regions, and establish communication channels between stationary points and mobile points in the analysis regions; Based on the communication channel, multiple mobile points from the same mobile location are guided to different analysis areas on the data panel. The driving point with the same name distributed in different analysis areas moves sequentially from the stationary point, and the data blocks in the data panel of its respective area are analyzed in parallel to obtain the analysis data. Extract and synthesize the sub-data blocks and their relationships obtained from all mobile points with the same name within their respective regions after analyzing all stationary points, and generate the data cluster corresponding to the analysis task objective.
[0010] Preferably, the step of driving the corresponding moving points distributed in different analysis areas to move sequentially along the stationary points, and performing parallel analysis on the data blocks in the data layout of their respective areas to obtain analysis data includes: Drive the moving point within each analysis region to move sequentially along the stationary points within its region; At each station, the mobile point analyzes the sub-data blocks carried by the data points associated with the station to obtain analysis results and determines whether the analysis results belong to the direct demand data of the current mobile point. If it belongs to the category, extract the sub-data blocks and their relationships that are directly related to the current analysis task objective as the analysis data for the current movement point; If it does not belong to the category, then when any moving point moves to the stationary point of an existing marked point, the enable judgment is performed.
[0011] Preferably, the step of performing the activation judgment when any moving point moves to a stationary point where an existing marker point already exists includes: For identified sub-data blocks or associations that are not directly related to the current target, a marker point is created at the station location for recording, and at least one applicable label is generated and bound to the marker point; When any mobile point moves to a station where an existing marker point is located, the mobile point's identity tag is matched with the applicable tag of the marker point; If the match is successful, the connection between the moving point and the marker point is established, and the sub-data blocks or relationships recorded by the marker point are directly obtained and integrated into the current analysis process. Once any mobile point successfully connects and uses the content stored in the marker point, the counter item corresponding to the current mobile point's identity tag is decremented. When the usage counter reaches zero, the marker point is automatically canceled and deleted.
[0012] Preferably, the step of inputting the data cluster into the trained analysis model, analyzing the data cluster based on the analysis model, and obtaining the text analysis result corresponding to the analysis task objective includes: Obtain the analysis models corresponding to the analysis task objectives after training is completed; obtain the target data clusters generated by analyzing multiple move points with the same name, as well as the analysis task objectives corresponding to the move points with the same name; The target data set is input into the analysis model corresponding to the analysis task objective for in-depth text analysis, and the text analysis results corresponding to the analysis task objective are obtained.
[0013] Compared with the prior art, the beneficial effects of the present invention are: A database is constructed using corresponding text data. Data clusters are extracted based on task requirements, and the corresponding text data within these clusters is input into a model tailored to the task requirements for analysis. This allows for flexible selection of data clusters that match the needs of the analysis task. Deep text analysis is then performed on the data clusters corresponding to the analysis task objective, adapting the analysis to different objectives and improving the accuracy of the analyzed text. Multiple identity tags for different analysis task objectives are stored at the movement points, and markers and their corresponding applicable tags are left at the corresponding locations for subsequent docking between the movement points and the markers. This increases the efficiency of data block analysis at the corresponding marker locations, thereby improving the overall efficiency of text analysis. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example: Please refer to Figure 1 This invention provides a technical solution for an intelligent natural language processing and deep text analysis method: an intelligent natural language processing and deep text analysis method, comprising the following steps: S1: Obtain the text data to be analyzed and construct a data layout corresponding to the text data to be analyzed; The steps for acquiring text data to be analyzed and constructing a data layout for the text data to be analyzed include: acquiring text data to be analyzed, dividing the text data into multiple data blocks, setting up anchor points for the multiple data blocks, establishing communication channels between the multiple anchor points in sequence, setting up multiple data areas for each anchor point, wherein each data area consists of multiple data points arranged in sequence, arranging multiple sub-data blocks in the data blocks in sequence on the data points, and sorting the multiple data points in sequence according to the division order of the data blocks to obtain the data layout; Specifically, the text data is divided into multiple data blocks, which can be a sentence or a paragraph. Multiple data blocks are laid out on corresponding data areas. Each data block contains multiple sub-data blocks, which are mapped one-to-one to data points. For example, when a data block is a paragraph, a sub-data block can be a word or a character. Each character is placed on its corresponding data point. All data blocks in the text data are laid out sequentially on the data points of the corresponding data areas according to the division order to obtain the corresponding text data layout. Stationary points are set in the database for multiple data blocks. Stationary points are used as the resting points for the movement points of the corresponding analysis task targets on the target disk on the data layout. The data blocks corresponding to the stationary points are analyzed at the stationary points. The entities and entity relationships after analysis are extracted or marked. For the marked entities and entity relationships, it is convenient for the movement points of other tasks to directly extract the corresponding entities and entity relationships when they reach the stationary point, so as not to repeat the analysis of the same data in the data block as the previous task, thus improving the efficiency of text data analysis.
[0018] S2: Obtain the analysis task objective of the text data, and construct a target disk based on the analysis task objective. The target disk consists of multiple moving points that correspond to different analysis task objectives. The steps for acquiring text data analysis task objectives and constructing a target disk based on these objectives include: acquiring multiple analysis task objectives corresponding to the same text data source; assigning a corresponding mobile point to each analysis task objective, with each mobile point storing multiple mobile points; configuring multiple marker points and multiple docking points within each mobile point; configuring multiple applicable tags for each marker point, where the applicable tags are identity tags of other mobile points that can be enabled by the content marked by the marker point; establishing communication connections between marker points and docking points, and setting trigger conditions, where the trigger condition is that the applicable tags corresponding to the marker point contain the identity tags of the mobile points corresponding to the docking points; and connecting the mobile points corresponding to multiple analysis task objectives to obtain the target disk. Specifically, each moving point is an independent data processing and analysis unit, used to execute its corresponding analysis task objective. It moves sequentially from stationary points on the data panel, analyzing the data blocks at each stationary point. Multiple marker points and docking points are set for each analysis task objective. Both docking points and marker points are stored in the moving point corresponding to the analysis task objective. Docking points on moving points corresponding to different analysis task objectives can dock with marker points corresponding to other analysis task objectives. Marker points are used to mark data points and the relationships between them. Marker points are configured with corresponding labels, which can include the data point and the relationships between it. The moving point corresponding to the analysis task objective is also defined. When a moving point corresponding to an analysis task objective analyzes data in a data block, it establishes relationships between data points based on the analysis process, extracts necessary relationships, and marks unnecessary relationships. When the corresponding association is marked, it is first determined whether the applicable label of the marker point is consistent with the identity label of the target mobile point. If they are consistent, it means that the triggering condition is met. The communication connection channel between the docking point and the marker point is temporarily opened to connect the marker point and the docking point. Then, through the docking point, the data points marked by the marker point and the association relationships between the data points can be directly obtained and enabled, saving the time of the current mobile point to re-analyze the marked data points and corresponding association relationships, thus improving the analysis efficiency of the current mobile point. When any mobile point performs an analysis task, the marker points of other mobile points are scanned through the docking point of that mobile point. Based on the matching result of the applicable label of the marker point and the identity of the current mobile point, and the relevance assessment of the content recorded by the marker point and the current analysis requirements, the docking is established automatically or after confirmation. Through the established docking channel, the data points and association relationships recorded by the marker point are directly imported into the analysis context of the current mobile point to realize cross-task knowledge use.
[0019] The steps of configuring multiple marker points within each mobile point and configuring multiple applicable tags for each marker point include: configuring a unique identity tag for each mobile point where the analysis task target is located, and storing multiple identity tags corresponding to different analysis task targets in each mobile point to obtain an identity tag library; obtaining the data point where the mobile point is currently located and its corresponding association as marker information, selecting multiple identity tags from the identity tag library based on the marker information, and temporarily assigning the selected identity tags as applicable tags for the corresponding current marker point; and setting a usage counter for each marker point, the initial value of which is determined according to the number of types of applicable tags configured for the marker point.
[0020] It should be noted that the identity tag of the mobile point where the task target is located is the same as the identity tag of the task target. Whenever a mobile point with a specific identity tag successfully connects to and uses the marker point, regardless of the number of times the mobile point has connected, only one decrement operation is performed on the counter, recording the correspondence between each applicable tag and one or more mobile point identity tags. The initial value of the usage counter is equal to the number of different mobile point identity tags that can connect to the marker point, as parsed from the tag mapping table. When a subsequent mobile point moves to the location of the marker point, its identity tag is matched with the applicable tag of the marker point. If the match is successful, it connects and enables the marker content. Whenever a marker point is successfully connected and used, its usage counter is decremented. When the usage counter of a marker point reaches zero, or when all mobile points have completed the analysis, the corresponding marker is automatically cleaned up. If a marker point is successfully enabled by multiple mobile points with different identity tags, each successful activation will cause the count item corresponding to the mobile point's identity tag in its usage counter to be decremented independently, protecting the timeliness and recycling of applicable tags, avoiding outdated or obsolete marker points from consuming resources ineffectively, thereby improving the system's operating efficiency. Specifically, at the current analysis location of the moving point, a marker point is created to record the identified data points and their relationships at that location; the dynamically allocated applicable tags are bound and stored with the marker point; the access permissions of the information recorded by the marker point are limited to the specific analysis task target range corresponding to the applicable tag; when any moving point moves to a location on the target disk, the main identity tag of the current moving point is obtained; it is checked whether there are marker points left by other moving points at that location; the main identity tag of the current moving point is matched with the applicable tags bound to each marker point at that location; if the match is successful, a connection is automatically or after authorization between the current moving point and the marker point is established. The method retrieves the information recorded by the marker point. The conditions for a successful match are: the primary identity label of the current mobile point is exactly the same as the applicable label of the marker point. The label is not preset, but is temporarily assigned based on the current data point location of the mobile point. Based on which other mobile points corresponding to the analysis task objectives may be applicable to the current data point location and its corresponding association, the identity label of the mobile point is assigned to the marker point, and the applicable label is temporarily assigned to make the information marked by the marker point more applicable to the corresponding mobile point. This increases the accuracy of the data point location and its corresponding association when performing text analysis at the corresponding mobile point, and at the same time improves the analysis efficiency of the mobile point at that point.
[0021] S3: Establish the relationship between the target data panel and the data layout, and extract the data clusters corresponding to the target points from the data layout based on the relationship; The steps for establishing the association between the target disk and the data layout, and extracting data clusters corresponding to the target points from the data layout based on the association, include: obtaining the mobile points on the target disk, and simultaneously deploying multiple identical mobile points at a single mobile point; dividing the data layout composed of the target text data into regions to form multiple analysis regions, and establishing communication channels between the stationary points and mobile points in the analysis regions; based on the communication channels, guiding multiple mobile points from the same mobile point to different analysis regions of the data layout; driving the same-named mobile points distributed in different analysis regions to move along the stationary points, and sequentially performing parallel analysis on the data blocks in the data layout of their respective regions to obtain analysis data; extracting and integrating the sub-data blocks and associations obtained by all the same-named mobile points in their respective regions after analysis of all stationary points, to generate data clusters corresponding to the analysis task objectives; It should be noted that the same movement point refers to a movement point corresponding to the same identity tag of the same analysis task target. A communication channel is established between the target disk and the data panel. Multiple movement points are first set at the movement point location. The movement points can randomly land on the first stationary point in the corresponding area. Movement points released from multiple different movement point locations can be analyzed simultaneously on the same stationary point. The data panel is divided into regions, and multiple movement points from the same movement point location are released to stationary points in different data regions for simultaneous analysis. The data analyzed by multiple movement points corresponding to the same movement point location are combined to obtain a data cluster. The data cluster contains text data from multiple data points collected after the analysis of multiple movement points with the same name, as well as the correlation between sub-data blocks. The data cluster is input into the analysis model corresponding to the analysis task target. Based on the data blocks and the correlation between data blocks (that is, the correlation between data points and data points, where data points refer to the data points where data blocks are located), the analysis model analyzes the text depth of the text data on the analysis task target to obtain the analysis results.
[0022] The steps for driving mobile points with the same name distributed in different analysis areas to move sequentially along the stationary points and sequentially perform parallel analysis on the data blocks in the data panel of their respective areas to obtain analysis data include: driving mobile points in each analysis area to move sequentially along the stationary points in their respective areas; at each stationary point, the mobile point analyzes the sub-data blocks carried by the data points associated with the stationary point to obtain analysis results, and determines whether the analysis results belong to the direct requirement data of the current mobile point. If they do, the sub-data blocks and their relationships directly related to the current analysis task objective are extracted as the analysis data of the current mobile point; if they do not, when any mobile point moves to a stationary point with an existing marker, an activation judgment is performed. When any mobile point moves to a station with an existing marker, the activation judgment steps include: for identified sub-data blocks or relationships not directly related to the current target, creating a marker at the station location for recording, and generating at least one applicable tag to bind to the marker; when any mobile point moves to a station with an existing marker, matching the mobile point's identity tag with the marker's applicable tag; if the match is successful, establishing a connection between the mobile point and the marker, directly obtaining the sub-data blocks or relationships recorded by the marker, and integrating them into the current analysis process; after any mobile point successfully connects and uses the content stored by the marker, the counter item corresponding to the current mobile point's identity tag is decremented; when the usage counter reaches zero, the marker is automatically canceled and deleted. It should be noted that the rules for generating applicable tags are as follows: based on the content characteristics, relationship types, and predefined knowledge graph of the current stationary sub-data block, the mobile point predicts which other analysis task objectives this information may be valuable to; based on the prediction results, the identity tag corresponding to the target analysis task is selected from the identity tag library stored by the current mobile point and temporarily assigned as the applicable tag for the marker point. Specifically, a communication channel is established between the stationary point and the mobile point. The mobile point moves sequentially through the communication channel, extracting the content it needs and marking the content it doesn't need. The marking points are automatically canceled based on the number of times they are used, with the number of uses corresponding to the types of applicable tags. Once the applicable tags are matched, the marking is automatically canceled. Communication connections are established between the marking points and the mobile points, and multiple mobile points can also communicate with each other. Once multiple mobile points have obtained the required data analysis text, the corresponding markings are also automatically canceled. The adaptation tags stored in the marking points are temporarily assigned by the mobile points at their current locations. The mobile points store multiple identity tags for different analysis task targets and leave marking points and the applicable tags corresponding to the marking information at the corresponding locations for subsequent docking between the mobile points and the marking points, thereby increasing the efficiency of the mobile points in analyzing data blocks at the corresponding marking point locations. Specifically, the target dashboard serves as the command and dispatch center for analytical tasks. It is a virtual or visual management plane that carries out the planning, deployment, and monitoring of all analytical tasks. It consists of multiple mobile points, each corresponding to a specific analytical task type. For example, in the analysis of the new energy vehicle industry, the target dashboard includes three mobile points: a technology analysis point, a supply chain analysis point, and a policy analysis point. Mobile points act as execution agents for analytical tasks. They are analytical units dispatched from the target dashboard, possessing specific identities and capabilities. They move along the data plane along the stationed points, analyzing and extracting data from the sub-data blocks covered by the stationed points, creating markers to retain useful information, and activating markers for other mobile points. For example, the mobile point "Tech_Analy"... "st_01", with the identity tag "Technical Analysis", is specifically used to extract battery technical parameters and performance data. A stationary point is a fixed analysis location node on the data layout, serving as a resting and analysis station for moving points. It identifies the location on the data layout and stores the marker points created at that location. It provides a stable analysis environment for moving points, hosts data blocks and sub-data blocks, manages the local marker point set, and coordinates concurrent access from multiple moving points. A data area is a logical partitioning unit on the data layout, a continuous region composed of multiple stationary points and related data. It is the basic unit for enabling parallel task execution, providing local data continuity guarantees, and supporting load balancing between areas. For example, the data area "Technical Chapter Area" contains all stationary points from Chapters 1-3. This section focuses on analyzing technical content. The data layout is a structured plane that carries the text data to be analyzed. It serves as the management interface for the entire analysis activity and can be a two-dimensional array in memory, a collection of tables in a database, or a directory structure in a file system. It physically carries all the analysis data, provides a spatialized analysis environment, manages data access permissions and concurrency, and records the analysis activities. Data blocks are the basic analytical units on the data layout, corresponding to a logical paragraph or semantically complete segment of text. Sub-data blocks are subdivided units of data blocks, representing the smallest data segments with independent semantics or functions. Data points are the precise coordinates on the data layout used to carry sub-data blocks. Marker points are created by the moving points during the analysis process and record valuable information. The knowledge anchors are created when a mobile point identifies valuable but not directly needed information, retaining the valuable information for later use. Tags enable knowledge reuse and avoid repeatedly analyzing the same content. Applicable tags are metadata tags for the tags, describing which analysis tasks the tag's content is suitable for reuse. This helps other mobile points quickly find the information they need, controls which tasks can access it, and manages effectiveness in conjunction with counters. The counter is the lifecycle management mechanism for the tags, recording the number of times a tag is reused and its status. Each applicable tag corresponds to an independent counter. After a mobile point is successfully reused, only the counter matching its identity is decremented. Multiple reuses of the same mobile point do not count repeatedly, preventing tags from occupying resources indefinitely.A data cluster is a structured knowledge set formed by fusing the results of multiple mobile points with the same name after analysis, used to support subsequent in-depth analysis. An identity tag is a unique identifier for each mobile point, determining its analysis task type and permissions, controlling access permissions, and guiding collaboration between mobile points. A communication channel is a data transmission path connecting the target disk and the data layout, used for mobile point deployment, communication connections between multiple stations, and movement of mobile points between stations, enabling collaborative work among multiple mobile points. It can quickly find suitable text data for the corresponding analysis task objective from text data, and then use the analysis model to perform in-depth analysis of the text data according to the corresponding analysis task objective, improving the accuracy of the analysis.
[0023] It should be noted that the analysis process between the mobile point and its associated sub-data blocks at the stationary point is as follows: When the mobile point reaches the target stationary point on the data page, the data block associated with that stationary point is obtained. Multi-granularity parsing is performed on the data block to obtain multiple candidate sub-data blocks and their semantic features. The analysis task objective corresponding to the current mobile point is transformed into a set of requirement feature vectors. These requirement feature vectors include: key entity types, types of relationships of interest, and attribute constraints. The semantic features of each candidate sub-data block are matched with the requirement feature vectors to obtain a matching score. Based on the matching score, classification processing is performed: if the matching score is higher than a preset threshold, the candidate sub-data block is determined to be direct requirement data and is extracted as a requirement sub-data block; if the matching score is lower than the preset threshold, the candidate sub-data block is determined to be potentially related data and is retained along with the extracted sub-data blocks. The relationship information is used for subsequent analysis; irrelevant data is not processed; for the extracted requirement sub-data blocks, relationship mining is performed: identifying entities appearing in the requirement sub-data blocks; identifying direct relationships between entities based on syntactic analysis and semantic role labeling; mining indirect relationships between entities based on co-occurrence statistics and contextual reasoning (the specific method for mining indirect relationships is: within the current data block, if two entities are not directly connected syntactically, but frequently co-occur in the same context window; and the type combination of the two entities belongs to a predefined set of associative entity pairs; then based on their simultaneous occurrence frequency, relative position, and shared modifiers, an indirect relationship with confidence is inferred and generated); the identified relationships are structured and stored according to their type and strength to form a set of association relationships; the analysis results of the current data block are output, including: a list of extracted requirement sub-data blocks and the corresponding set of association relationships; The multi-granularity parsing of the data block specifically includes: First granularity: dividing the data block into sentence-level sub-data blocks according to punctuation marks and conjunctions; Second granularity: performing word segmentation and part-of-speech tagging on each sentence-level sub-data block to obtain vocabulary-level sub-data blocks; Third granularity: performing dependency parsing on the sentences to identify phrase-level sub-data blocks centered on core predicates; Semantic features are extracted for each sub-data block at each granularity, including entity type distribution, keyword vectors, sentiment polarity, and topic tags.
[0024] In the process of analyzing text data through marker points, in addition to analyzing whether the sub-data block is applicable to the current analysis task objective, it is also necessary to analyze which analysis task objective the content corresponding to the sub-data block tends to. Based on the applicable analysis task objective, the marker point is marked and assigned a corresponding identity label. As the moving point moves in sequence along the stationary point, it only analyzes which analysis task objective the sub-data block covered by the stationary point is useful to, and directly extracts the sub-data block that is useful to itself. For the sub-data block that is not useful to itself, it is marked. This is equivalent to the initial data filtering and extraction, extracting the text data that corresponds to the needs of the analysis task objective from the multi-source data. The unused sub-data blocks are only marked and not extracted. In addition, during the analysis process, it is determined that the sub-data blocks that are not used by the current analysis task objective may be applicable to other analysis task objectives, and assigns corresponding identity labels to them. When the moving point corresponding to other analysis task objectives moves to that position, the corresponding sub-data block is directly extracted through the identity label, without having to repeat the analysis of the sub-data block corresponding to the marker point. This improves the efficiency of the moving point in filtering and extracting text data, and thus improves the efficiency of text analysis.
[0025] S4: Input the data cluster into the trained analysis model, analyze the data cluster based on the analysis model, and obtain the text analysis results corresponding to the analysis task objectives; The steps of inputting data clusters into a trained analysis model and analyzing the data clusters based on the analysis model to obtain the text analysis results corresponding to the analysis task objectives include: obtaining the analysis models corresponding to the trained analysis task objectives; obtaining the target data clusters generated by analyzing multiple homonymous moving points and the analysis task objectives corresponding to the homonymous moving points; inputting the target data clusters into the analysis model corresponding to the analysis task objectives to perform deep text analysis and obtain the text analysis results corresponding to the analysis task objectives.
[0026] Same-name movepoints refer to movepoints that correspond to the same analytical task objective but are scattered across different data areas.
[0027] Specifically, the analysis model is trained based on historical data clusters and corresponding expert annotation results, and has the ability to process structured data cluster inputs. The target data cluster contains the following structured information: multiple extracted named entities and their attributes, a relationship graph between entities (the relationship graph includes relationship type and strength weight), and analysis metadata (including data source distribution, timestamp sequence, and regional contribution statistics). The target data cluster is preprocessed and adapted, converting the heterogeneous data structure of the data cluster into a unified input format required by the analysis model, and mapping the entities and relationship terms in the data cluster to the semantic space of the analysis model. The adapted data cluster is then input into the analysis model to perform multi-level collaborative analysis: the first level of analysis: entity-level pattern recognition, identifying... The analysis process involves four layers: first, layer 1: 1) identifying the distribution patterns, clustering characteristics, and outliers of entities within data clusters; second, layer 2: relational network reasoning, performing path analysis, community discovery, and influence calculation based on relational networks; third, layer 3: cross-regional association mining, utilizing regional contribution data from metadata to analyze association patterns and dependencies between different data regions; and fourth, layer 4: temporal evolution inference, analyzing the dynamic trends of entities and relationships based on timestamp sequences. The output is structured text analysis results, including: core findings and conclusions related to the analysis task objectives, supporting data item citations and statistical evidence, confidence scores and uncertainty ranges for the conclusions, specific operational suggestions based on the analysis results, and incremental update suggestions for the knowledge base of the analysis model. A database is constructed corresponding to the text data. Data clusters are extracted according to task requirements, and the corresponding text data within these clusters is input into the model corresponding to the task requirements for analysis. This allows for flexible selection of data clusters corresponding to the analysis task requirements, adapting to different analysis objectives and improving the accuracy of the analyzed text. In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0028] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An intelligent natural language processing and deep text analysis method, characterized in that, Includes the following steps: Obtain the text data to be analyzed, and construct a data layout corresponding to the text data to be analyzed; The analysis task objective is to acquire text data. Based on the analysis task objective, a target disk is constructed. The target disk consists of multiple moving points, each corresponding to a different analysis task objective. Establish the relationship between the target data panel and the data layout, and extract the data clusters corresponding to the target points from the data layout based on the relationship; The data clusters are input into the trained analysis model, and the data clusters are analyzed based on the analysis model to obtain the text analysis results corresponding to the analysis task objectives.
2. The intelligent natural language processing and deep text analysis method according to claim 1, characterized in that: The steps of obtaining the text data to be analyzed and constructing a data layout corresponding to the text data to be analyzed include: Obtain the text data to be analyzed, divide the text data into multiple data blocks, and set station points for the corresponding multiple data blocks; Communication channels are established sequentially between multiple outposts, and multiple data areas are set for each outpost. Each data area consists of multiple data points arranged sequentially. Multiple sub-data blocks within a data block are arranged sequentially on data points, and these data points are then sorted according to the order in which the data blocks are divided to obtain the data layout.
3. The intelligent natural language processing and deep text analysis method according to claim 1, characterized in that: The steps for constructing a target disk based on the analysis task objective of acquiring text data include: Get multiple analysis task targets corresponding to the same text data source, set a corresponding moving point for each analysis task target, and store multiple moving points in each moving point; Multiple marker points and multiple docking points are configured within each mobile point. Multiple applicable tags are configured for each marker point. The applicable tags are the identity tags of other mobile points that can be enabled by the content marked by the marker point. Establish a communication connection between the marker point and the docking point, and set the trigger condition, wherein the trigger condition is that the applicable tag corresponding to the marker point contains the identity tag of the mobile point corresponding to the docking point; The target disk is obtained by connecting the movement points corresponding to multiple analysis task objectives.
4. The intelligent natural language processing and deep text analysis method according to claim 3, characterized in that: The step of configuring multiple marker points within each movement point and configuring multiple applicable tags for each marker point includes: Configure a unique identity tag for each mobile point where the analysis task target is located, and store multiple identity tags corresponding to different analysis task targets at each mobile point to obtain an identity tag library; The current location of the mobile point and its corresponding relationship are obtained as the tagging information. Based on the tagging information, multiple identity tags are selected from the identity tag library, and the selected identity tags are temporarily assigned as the applicable tags for the current tagging point. A counter is set for each marker point, and the initial value of the counter is determined based on the number of applicable tag types configured for the marker point.
5. The intelligent natural language processing and deep text analysis method according to claim 1, characterized in that: The step of establishing the association between the target disk and the data layout, and extracting the data clusters corresponding to the target points from the data layout based on the association, includes: Obtain the mobile points on the target disk, and deploy multiple identical mobile points simultaneously on a single mobile point; divide the data layout composed of target text data into regions to form multiple analysis regions, and establish communication channels between the stationary points and mobile points in the analysis regions; Based on the communication channel, multiple mobile points from the same mobile location are guided to different analysis areas on the data panel. The driving point with the same name distributed in different analysis areas moves sequentially from the stationary point, and the data blocks in the data panel of its respective area are analyzed in parallel to obtain the analysis data. Extract and synthesize the sub-data blocks and their relationships obtained from all mobile points with the same name within their respective regions after analyzing all stationary points, and generate the data cluster corresponding to the analysis task objective.
6. The intelligent natural language processing and deep text analysis method according to claim 5, characterized in that: The steps of driving the corresponding moving points distributed in different analysis areas to move sequentially along the stationary points, and sequentially performing parallel analysis on the data blocks in the data layout of their respective areas to obtain the analysis data include: Drive the moving point within each analysis region to move sequentially along the stationary points within its region; At each station, the mobile point analyzes the sub-data blocks carried by the data points associated with the station to obtain analysis results and determines whether the analysis results belong to the direct demand data of the current mobile point. If it belongs to the category, extract the sub-data blocks and their relationships that are directly related to the current analysis task objective as the analysis data for the current movement point; If it does not belong to the category, then when any moving point moves to the stationary point of an existing marked point, the enable judgment is performed.
7. The intelligent natural language processing and deep text analysis method according to claim 6, characterized in that: The step of performing the activation judgment when any moving point moves to the stationary point of an existing marked point includes: For identified sub-data blocks or associations that are not directly related to the current target, a marker point is created at the station location for recording, and at least one applicable label is generated and bound to the marker point; When any mobile point moves to a station where an existing marker point is located, the mobile point's identity tag is matched with the applicable tag of the marker point; If the match is successful, the connection between the moving point and the marker point is established, and the sub-data blocks or relationships recorded by the marker point are directly obtained and integrated into the current analysis process. Once any mobile point successfully connects and uses the content stored in the marker point, the counter item corresponding to the current mobile point's identity tag is decremented. When the usage counter reaches zero, the marker point is automatically canceled and deleted.
8. The intelligent natural language processing and deep text analysis method according to claim 1, characterized in that: The steps of inputting the data cluster into the trained analysis model, analyzing the data cluster based on the analysis model, and obtaining the text analysis results corresponding to the analysis task objective include: Obtain the analysis models corresponding to the analysis task objectives after training is completed; obtain the target data clusters generated by analyzing multiple move points with the same name, as well as the analysis task objectives corresponding to the move points with the same name; The target data set is input into the analysis model corresponding to the analysis task objective for in-depth text analysis, and the text analysis results corresponding to the analysis task objective are obtained.