Big data information analysis method and system based on artificial intelligence

By identifying the direction vector and channel number of the jump structure in the trajectory, splitting the path segments, and analyzing the bridge connection relationship, the identification problem at the path direction mutation points is solved, the direction continuity of the path sequence and the stability of the semantic chain are achieved, and the accuracy of information analysis is improved.

CN120541546AActive Publication Date: 2025-08-26FUYING TECHNOLOGY (JIAXING) CO LTD +1

Patent Information

Application Number
CN202510637417.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-26
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing artificial intelligence big data information analysis method is difficult to achieve differentiated segment jump recognition based on connection characteristics at the mutation points in the path direction, and lacks directional continuity of path sequence, resulting in the inability to identify the break points of the semantic chain and the lack of actual connection conditions for the migration of bridge segments between paths, which affects the integration accuracy of structural output.

Method used

By obtaining the direction vector of the jump structure in the trajectory, combining the channel number and path number, identifying the direction mutation points, splitting the path segments that do not have upstream and downstream mapping relationships, analyzing the bridge segment connection relationships, reordering the path segments to maintain direction consistency, and constructing a semantic behavior focus analysis set.

Benefits of technology

It improves the accuracy of node sequence structure positioning, maintains path continuity, enhances the consistency of path recognition in the channel aggregation area and the stability of node attributes, and constructs a behavioral aggregation sequence of semantic chains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541546A_ABST
    Figure CN120541546A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a big data information analysis method and system based on artificial intelligence, and the method comprises the following steps: obtaining a jump segment direction vector, classifying label nodes with consistent directions, splitting a path lacking a mapping relation to form an interruption segment, recognizing a node with continuous directions as a bridge segment entrance, and obtaining a bridge segment direction vector; according to the method, by means of classification according to the node sequence and the channel number in the segment hopping direction, direction aggregation boundary determination and semantic field combination path position interruption recognition, the node sequence structure positioning precision and the paragraph division reasonability are improved, and the focus analysis set is generated by combining the path segments. Bridge segment connection is screened based on direction consistency, path continuation logic, behavior tracks and aggregation numbers are kept in sequence and hierarchy regularization, a behavior aggregation sequence of a semantic chain is constructed, and the continuity of path recognition in a channel aggregation area and the stability of node attribution are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a big data information analysis method and system based on artificial intelligence. Background Art

[0002] The field of artificial intelligence (AI) encompasses the ability to intelligently reason about complex problems, understand data, and make behavioral decisions based on machine learning, deep learning, natural language processing, knowledge graphs, neural network modeling, and expert systems. The core of this technology lies in the use of adaptive and generalizable mathematical models to automatically learn and infer structured or unstructured data, thereby forming an information processing process with the capabilities of perception, understanding, and feedback. AI systems are widely used in speech recognition, image recognition, intelligent recommendations, intelligent search, and intelligent decision support, among other areas. Key processes include sample training, model building, parameter optimization, data input parsing, and output response control.

[0003] Among them, the AI-based big data information analysis method refers to the use of neural network learning mechanisms and graph neural structure modeling mechanisms, based on a unified feature vector set constructed from structured log data and semi-structured text data, to perform dimensionality compression and local feature extraction on multidimensional data through convolution transformation functions. Combined with a feature focusing strategy based on an attention mechanism, it performs key indicator clustering and logical sequence pattern analysis on the input data, and completes the inductive classification of abnormal behaviors and potential patterns through probabilistic inference methods. Based on the detailed records and time series in the historical data set, this solution uses a graph construction mechanism to generate a node attribute relationship matrix, and generates a result label set based on the construction principles of a cluster classification tree.

[0004] Existing AI-based big data information analysis methods rely on feature vector construction to globally model nodes. This makes it difficult to differentiate hops based on connection features at path directional abrupt changes, and path ordering cannot preserve directional continuity in multi-directional intersection scenarios. Regarding semantic field processing, multi-dependency models employ abstract latent representations and lack an explicit synchronization mechanism for the order of semantic fields within node segments, resulting in the inability to identify semantic chain breaks in some behavioral paths. Inter-path bridge migration does not rely on the actual connectivity conditions between hops for structural extension. Connectivity predictions are often based on node similarity or graph adjacency probabilities, lacking logical judgment of connection direction, resulting in a lack of directional continuity in path relay construction. Trajectory numbers and cluster numbers are classified solely based on label feature matching, without any correlation between their positional order and hierarchical proportion in the path structure. This can easily lead to classification distortions where semantically identical trajectories are misaligned. Within channel aggregation paths, the lack of a behavioral chain regularization method based on aggregation hierarchy makes it difficult to avoid semantic redundancy and node hop overlap during path behavior merging, impacting the accuracy of the integrated structural output. Summary of the Invention

[0005] In order to solve the technical problems existing in the prior art, an embodiment of the present invention provides a big data information analysis method based on artificial intelligence, comprising the following steps:

[0006] In order to achieve the above objectives, the present invention adopts the following technical solution: a big data information analysis method based on artificial intelligence, comprising the following steps:

[0007] S1: Obtain the jump point direction vector for jump structure analysis in the trajectory, compare the channel number of the direction mutation point with the path number and the order of adjacent jump points, merge the label nodes according to the direction consistency, and obtain the direction jump aggregation label group map;

[0008] S2: Based on the identified node path in the direction jump aggregation label group graph, according to the label node in the aggregation segment, read the semantic track field and path number, split the path segment without upstream and downstream mapping relationship, and obtain the semantic interrupt structure mapping fragment set;

[0009] S3: Based on the break nodes in the semantic interruption structure mapping fragment set, the connection numbers and direction sequences of the forward and backward jump segments are analyzed, and the discontinuous segments with consistent directions are marked as bridge segment entrances to obtain a bridge segment entrance path node structure diagram;

[0010] S4: Based on the bridge segment connection relationship in the bridge segment entrance path node structure diagram, the jump direction sequence is compared, and the path segments with consistent offsets and continuous levels are reordered to obtain a path slip structure connection sequence diagram;

[0011] S5: Based on the label node group in the path slip structure connected sequence diagram, the semantic extension structure in the channel aggregation area is analyzed, and the continuous trajectory and semantic distribution are merged into behavior segments to obtain a semantic behavior focus analysis set.

[0012] As a further solution of the present invention, the direction jump aggregation label group map includes a jump point direction vector sequence, a channel number sequence mapping set, a direction convergence label node group, and a path structure branch paragraph set; the semantic interruption structure mapping fragment set includes a path paragraph number sequence, a semantic field broken link node set, a discontinuous mapping path index, and a structure interruption node classification result; the bridge section entry path node structure diagram includes a jump section connection number index, a path direction sequence set, a connectable path fragment set, and a bridge section entry node identification set; the path slip structure connection sequence diagram includes a jump direction offset sorting group, a structure hierarchy continuous node group, a path linkage slip segment number sequence, and a channel sequence mapping table; the semantic behavior focus analysis set includes a behavior trajectory aggregation sequence, a cluster number classification table, a path hierarchy relationship set, and a semantic pattern structure item.

[0013] As a further solution of the present invention, the specific steps of S1 are:

[0014] S101: Obtain a node set of the jump structure in the trajectory, extract the direction vector of the node and the connectivity sequence between adjacent nodes, combine the arrangement trend of the direction vector with the node jump position sequence, analyze the direction structure differences of the label nodes in the direction mutation area, and obtain the jump point direction offset relationship;

[0015] S102: Based on the jump point direction offset relationship, the channel number and path number of the label node where the mutation point is located are extracted, the channel number arrangement order and the number position of the jump point in the path are combined, and a channel node sequence with consistent direction offset trend is selected and sequentially integrated to obtain a channel label sequence structure;

[0016] S103: Based on the channel label sequence structure, the path numbers and channel structure sequences to which the label nodes belong are segmented by number, label combinations with similar directions in the same path are collected and the path structure is split, the obtained label node sequences are branched and clustered to obtain a directional jump aggregation label group map.

[0017] As a further solution of the present invention, the specific steps of S2 are:

[0018] S201: Based on the node path identified in the direction jump aggregation tag group graph, extract the tag node number in the aggregation segment, read the corresponding semantic track annotation field and path segment number, and pair and associate the semantic fields in the order of the path number to obtain a semantic field path mapping relationship;

[0019] S202: calling the semantic field path mapping relationship, screening the path number sequence that does not correspond to the semantic field continuously, identifying the label nodes where the semantic correspondence relationship is interrupted, extracting the identification values ​​of the interrupted nodes in the order of the path numbers and classifying and arranging them, to obtain a mapping structure interrupted node sequence;

[0020] S203: Based on the mapping structure interruption node sequence, the path numbers and paragraph positions of the included nodes are extracted, the label path where the interruption point is located is segmented according to the path number, and the corresponding path paragraph numbers and interruption node numbers are matched and merged to obtain a semantic interruption structure mapping fragment set.

[0021] As a further solution of the present invention, the specific steps of S3 are:

[0022] S301: Based on the path break nodes in the semantic break structure mapping fragment set, extract the preceding and following hop segment connection numbers and path direction sequence corresponding to the nodes, organize the hop segment connection structure according to the path numbers, sort the hop segment numbers in sequence and compare them with the path directions to obtain the path hop segment direction sequence;

[0023] S302: Calling the path hop direction sorting method, extracting the order of adjacent hops in the path direction sequence, calculating the frequency of repeated occurrence of the path direction in the hop sequence, marking the hop node number combinations with the same repetition frequency, and obtaining the overlap position of the direction sequence;

[0024] S303: Based on the overlapping positions of the direction sequence, node pairs with direction continuity characteristics are screened, path numbers and entry node positions are extracted, and the jump segment path numbers are classified and integrated according to the access sequence to obtain a bridge segment entry path node structure diagram.

[0025] As a further solution of the present invention, the calculation formula for the repetition frequency of the path direction in the hop sequence is specifically:

[0026]

[0027] Among them, F rep represents the frequency of repetition of the path direction in the hop sequence, D z Represents the cumulative number of occurrences of path direction z in the hop sequence, represents the mean number of occurrences of the path direction, G z represents the mean square sum of the differences between the position numbers of adjacent hops in the path in direction z, V z represents the number of path segments corresponding to direction z in the hop sequence, R z The distribution range of the path segments in direction z covers the number of hops, L z represents the number of hop label nodes in the path segment where direction z is located, and u represents the total number of path directions in the hop sequence.

[0028] As a further solution of the present invention, the specific steps of S4 are:

[0029] S401: Based on the bridge segment connection nodes in the bridge segment entrance path node structure diagram, extract the jump sequence between adjacent nodes, collect the start and end node numbers and jump directions of the path segment, and combine them in the order of path numbers to obtain a path direction sequence combination;

[0030] S402: Calling the path direction sequence combination, extracting the direction identifier and structure level value of the path segment, calculating the frequency of occurrence of the direction identifier in the path sequence, screening the path numbers with the same frequency and the same structure level, and obtaining the direction structure frequency combination path;

[0031] S403: Based on the directional structure frequency combined path, extract the jump node number sequence corresponding to the classified path segment, align the node number and path number position according to the directional frequency classification order, and obtain a path slip structure connection sequence diagram.

[0032] As a further solution of the present invention, the calculation formula for the frequency of occurrence of the direction identifier in the path sequence is specifically:

[0033]

[0034] in, Represents the frequency of occurrence of the i-th type direction mark in the path sequence, n i represents the total number of occurrences of the i-th type direction identifier in the current path direction sequence, w ij represents the direction weight value of the i-th type direction mark in the j-th occurrence, l ij Represents the structural level value of the path segment where the i-th type direction identifier appears for the jth time, m ij represents the path occupancy density of the structural unit connected by the path segment where the direction mark of type i appears for the jth time, s ij Represents the number of adjacent direction identifiers connected to the i-th type direction identifier at the j-th path segment position.

[0035] As a further solution of the present invention, the specific steps of S5 are:

[0036] S501: Based on the label node group in the path slip structure connection sequence diagram, extract the corresponding channel aggregation segment number, collect semantic field information of the label node in the aggregation segment, identify the node sequence with path extension relationship between the fields, merge the paths according to structural continuity, and obtain the semantic extension structure node;

[0037] S502: Calling the semantic extension structure node, extracting the behavior trajectory number, cluster number, and aggregation level value of the associated node in the path segment, analyzing the corresponding structural features between the trajectory number and the cluster number, classifying and integrating the label nodes whose number relationships have been verified, and obtaining the behavior trajectory cluster association sequence value;

[0038] S503: Based on the behavior trajectory cluster association sequence value, extract the label node number group with consistent structural hierarchy in the path segment, combine and associate the behavior trajectory with the node number, regularize the node behavior with association characteristics in the channel aggregation area, and obtain the semantic behavior focus analysis set.

[0039] A big data information analysis system based on artificial intelligence, comprising:

[0040] The direction cluster identification module obtains the direction vector of the jump structure node, extracts the channel number and path number of the mutation point label node, combines the direction vector sequence with the channel arrangement position, combines the nodes with continuous directions, aggregates the path numbers according to the jump point order, and obtains the direction jump aggregation label group map;

[0041] The semantic mapping link breaking module jumps the node path in the aggregation tag group graph based on the direction, reads the semantic annotation field and path segment number of the tag in the aggregation segment, merges the path numbers of the nodes with discontinuous upstream and downstream semantic fields, and obtains a semantic interruption structure mapping fragment set;

[0042] The path connection and reorganization module extracts the jump segment connection number and path direction based on the broken nodes in the semantic interruption structure mapping fragment set, compares the connection node position and direction sequence, groups the jump segments with structural connection, and obtains the bridge segment entrance path node structure diagram;

[0043] The sliding structure linkage module extracts the jump sequence to form a path direction sequence based on the connection nodes in the bridge section entrance path node structure diagram, calls the path offset trend and the structure level value, combines the path segments with consistent structures, constructs a continuous jump path, and obtains the path sliding structure connection sequence diagram;

[0044] The behavior feature focusing module connects the label nodes in the sequence diagram based on the path slip structure, identifies the semantic extension structure of the channel aggregation segment, calls the behavior trajectory field, cluster relationship and aggregation level, summarizes the node groups with overlapping clusters and consistent structures, and obtains the semantic behavior focusing analysis set.

[0045] Compared with the prior art, the advantages and positive effects of the present invention are:

[0046] In the present invention, the direction aggregation boundary is determined by classifying the jump direction according to the node order and channel number, the semantic field is combined with the path position to identify the interruption, the node sequence structure positioning accuracy and the paragraph division rationality are improved, the bridge connection is based on the direction consistency screening, the path continuation logic is maintained, the behavior trajectory and the aggregation number are regularized in sequence and hierarchy, and the behavior aggregation sequence of the semantic chain is constructed, which enhances the coherence of path identification and the stability of node attribution in the channel aggregation area. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0048] Figure 1 Schematic diagram of the steps of the present invention;

[0049] Figure 2 It is a system module diagram of the present invention. DETAILED DESCRIPTION

[0050] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0051] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0052] In the embodiments of the present invention, the terms "image" and "picture" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same. The terms "of," "corresponding," and "corresponding" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same.

[0053] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0054] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0055] See also Figure 1 , an embodiment of the present invention provides a big data information analysis method based on artificial intelligence, comprising the following steps:

[0056] S1: Obtain the jump structure in the trajectory and extract the jump point direction vector. Extract the channel number and path number of the label node where the path direction mutation point is located. Sequentially associate them based on the jump point arrangement between channels. Group and integrate label nodes with similar direction between adjacent channels. Divide the label sequence in the aggregation result into multiple branch sections according to the path structure to obtain the directional jump aggregation label group map.

[0057] S2: Based on the direction jump, aggregate the identified node paths in the label group graph, read the semantic trajectory annotation fields and continuous path segment numbers of all label nodes in the aggregation segment, segment the label paths where the semantic fields do not have continuous mapping between upstream and downstream nodes, and classify the node paths at the mapping interruption into a broken link structure sequence to obtain a semantic interruption structure mapping fragment set;

[0058] S3: Mapping the path break nodes in the segment set based on the semantic discontinuity structure, extracting the connection node numbers and path direction order of the preceding and following segments of the discontinuity segment, comparing the position of the connection structure between the path segments, and screening the discontinuity segments with connection characteristics based on the path direction. Path segments that meet the connection conditions are grouped into the bridge segment entrance set to obtain the bridge segment entrance path node structure diagram;

[0059] S4: Based on all bridge segment connection nodes in the bridge segment entrance path node structure diagram, the jump sequence between adjacent nodes is extracted to form a path direction sequence. The label paths with the same direction deviation trend in the sequence are sorted, and the path segments with the same deviation direction and consistent structural hierarchy are recombined into linkage sliding structure paths to obtain the path sliding structure connection sequence diagram;

[0060] S5: Based on the path slip structure connecting each label node group in the sequence diagram, the semantic extension structure in the channel aggregation segment is identified, the behavior trajectory, clustering relationship and path aggregation level of the label nodes in the aggregation area are integrated and summarized, and the label behavior patterns with analysis value are aggregated into readable structure items to obtain a semantic behavior focused analysis set.

[0061] The directional jump aggregation label group map includes the jump point direction vector sequence, the channel number sequence mapping set, the direction convergence label node group, and the path structure branch paragraph set. The semantic interruption structure mapping fragment set includes the path paragraph number sequence, the semantic field broken link node set, the discontinuous mapping path index, and the structure interruption node classification result. The bridge section entry path node structure diagram includes the jump section connection number index, the path direction sequence set, the connectable path fragment set, and the bridge section entry node identification set. The path slip structure connection sequence diagram includes the jump direction offset sorting group, the structure hierarchy continuous node group, the path linkage slip segment number sequence, and the channel sequence mapping table. The semantic behavior focus analysis set includes the behavior trajectory aggregation sequence, the cluster number classification table, the path hierarchy relationship set, and the semantic pattern structure item.

[0062] The specific steps of S1 are:

[0063] S101: Obtain a node set of the jump structure in the trajectory, extract the direction vector of the node and the connectivity sequence between adjacent nodes, combine the arrangement trend of the direction vector with the node jump position sequence, analyze the direction structure differences of the label nodes in the direction mutation area, and obtain the jump point direction offset relationship;

[0064] Obtain the spatial position coordinates and time label values ​​of each node in the jump trajectory, arrange the nodes in chronological order, record the position change vector between each two adjacent nodes, record the coordinate difference vector between the current node and the successor node as the direction vector, then sort the multiple direction vectors, label the node set according to the angular trend of the direction change, obtain the time difference sequence of the jump points between the nodes, mark the nodes with a jump frequency of more than 3 times in a unit time period as frequent jump nodes, combine the direction vector turning angle of the frequent jump nodes, identify the jump segment positions with a direction change amplitude of more than 45 degrees, then obtain the channel number and path number of the labels where these jump segments are located, form the channel index set where the jump point is located in the trajectory, combine the arrangement order of the previous and next jump points in the label path number, record the time order of the appearance of the label in multiple channel paths, and arrange the angle formed by the direction vector turning position from large to small. At the same time, perform difference analysis on the angle change trend of the direction vector before and after the jump segment according to the time series, and calculate the position where the direction change sequence has a trend reversal in different time periods. Nodes are set, and groups of jump points with continuously rising or falling directions are identified. The angle difference between adjacent jump points is determined to see if it reaches a reference range of [30 degrees, 90 degrees]. If the angle difference before and after the jump point falls within this range, it is classified as a sudden change region. Then, based on the jump sequence numbers of these nodes in each channel, all segments with sudden changes in direction are found in the overall path. The angle offset values ​​between the jump point locations in each channel and the path trunk are compared. If the angle offset value is greater than the set direction offset judgment reference value of 45 degrees, the node is marked as a direction offset node. The preceding and following jump segment labels of each group of direction offset nodes in the path graph are read to see if a continuous path mapping relationship is formed in their jump sequence. If a jump segment path skips an intermediate sequence segment and directly connects, it is recorded as a jump segment path jump and mapped to each label index according to the jump segment number. Finally, the direction change nodes that meet the above conditions are marked as sudden change jump point groups along with the path number, jump sequence, and angle offset value, thus obtaining the jump point direction offset relationship.

[0065] S102: Based on the jump point direction offset relationship, the channel number and path number of the label node where the mutation point is located are extracted. The channel number arrangement order and the jump point number position in the path are combined to select the channel node sequence with the consistent direction offset trend and sequentially integrate them to obtain the channel label sequence structure.

[0066] Read the label node information corresponding to each mutation point from the offset node set, use the channel number and path number of the label node as the basic index field, and build a sequential list based on the relative position number of the node in the path sequence where the label jump point is located. The order of the channel number is based on the path timestamp sequence, and the labels belonging to the jump point are sorted in chronological order. If there are multiple label nodes corresponding to the same channel number in different paths, they are rearranged in ascending order of the path number. The direction offset vectors between the jump segments are analyzed, and the angle change value of the direction vector is read in each group of consecutive jump segments. If the positive and negative signs of the adjacent angle change values ​​are consistent, the group of jump points is defined as having the same direction offset trend. The screening method is to perform a differential operation on the jump segment direction difference sequence, record the number of consecutive segments with positive or negative difference changes, and if there are more than two consecutive segments and the angle change amplitude does not reverse the sign, it is marked as a segment with consistent direction trend. The angle change difference is set as the judgment basis, and the angle difference threshold range [20 degrees, 60 degrees] is used as the screening standard. If the hop direction difference is always within this interval, the directional change of this segment is considered to be stable and consistent. On this basis, all hops that meet the directional trend consistency relationship are merged in channel number order, and the start and end label numbers of each segment and the original path number are marked. The label node position order between each hop in the path is then integrated to ensure that the position number of each label group in the channel path remains monotonically increasing. If the hop number is reversed or the hop segment is out of bounds, the node combination of this segment is excluded from the integration. For example, if labels T1 to T5 belong to channel C7, the path numbers are P21, P22, P23, P24, and P25, the directional offset angles are 35°, 40°, 38°, and 36°, the directional trend does not change, and the numbers are arranged in chronological order, then T1-T5 are combined into a channel segment with consistent trend. The above process is repeated for multiple such combinations. Finally, all label node sets that pass the directional trend screening are connected in number order to form a stable and continuous label sequence chain, generating a channel label sequence structure.

[0067] S103: Based on the channel label sequence structure, the path numbers and channel structure sequences to which the label nodes belong are segmented, label combinations with similar directions within the same path are collected and the path structure is split, the obtained label node sequences are branched and clustered to obtain a directional jump aggregate label group map;

[0068] The path number and the channel number are combined to form a double-layer path index group. The label nodes in the index group are preliminarily segmented according to the path number from small to large. In the label sequence within each segment, the jump direction vector is called and the direction difference between adjacent nodes is analyzed. The node sequence with a difference in the range of 20 degrees to 50 degrees and the same direction is regarded as a directional convergence segment. Then, the consistency of the label channel number and the path number is checked to mark the label combination with the same path number and the same direction. A sequence record table is generated for each combination. The table records the channel, path and position index of the label in the combination. Then, combined with the sequence record table, the adjacent jump segments in the path with a sudden change of direction exceeding 50 degrees are split. These mutation nodes are used as path structure demarcation points. Independent path segment numbers are generated for the path segments before and after each demarcation point, and they are divided into structurally independent data segments. After the segmentation is completed, the directional density value of each segment after the path split is calculated. The directional density value is obtained by the ratio of the sum of all jump angles in the path segment to the number of jump segments, and the directional density value is used as a reference for the label node branch division. The segments with directional density values ​​between 30 and 40 degrees are marked as candidate segments for the same-direction clustering. The index numbers of the label nodes in multiple candidate segments are grouped, and the candidate segments with adjacent path numbers and consistent directions are merged to form a label direction aggregation candidate set. The candidate set is then rearranged according to the order of the labels in the original trajectory. The rearranged label groups are divided according to the intersection of the channel numbers. If two label groups have overlapping path numbers, consistent directions, and the number of intersection labels is not less than 3, they are determined to belong to the same cluster set. Finally, all label sets that meet the above conditions are output as the directional cluster boundaries of the jump path to obtain a directional jump aggregation label group map.

[0069] The specific steps of S2 are:

[0070] S201: Based on the node path identified in the direction jump aggregation label group graph, extract the label node number in the aggregation segment, read the corresponding semantic track annotation field and path segment number, and pair and associate the semantic fields in the order of the path number to obtain the semantic field path mapping relationship;

[0071] First, read the start and end node numbers of each aggregation segment, write the number sequence into the index list, perform path positioning operations on each number item and synchronously call its label number and corresponding channel number, then read the semantic trajectory annotation field from the semantic information table corresponding to the label number. This field represents the behavioral meaning of the label node in different stages. The semantic field is recorded in the form of scalar coding, such as state segment labels such as "behavior start", "state maintenance", and "event alternation". The record form is a comparison set of behavioral state codes and corresponding time periods. For example, label T024 has the annotation sequence [(B1, 0–15), (B2, 15–35)], which means that frames 0 to 15 are labeled as behavioral state B1, and frames 15 to 35 are labeled as B2. Read the semantic trajectory field of each label node in turn, and sort the path numbers in the same channel in ascending order. All semantic fields are sequentially paired according to the node's path number position in the channel. In the process, it is necessary to determine whether the path number of the label node is continuously increasing. If there is a jump, the path is reset. Vacant positions are left blank and marked as path breakpoints. Semantic fields before and after the breakpoints are mapped separately. Then, a mapping relationship is established for the semantic fields in all consecutive path segments, using the label number as the key. The mapping conditions are that the label numbers must be consistent in the order of the path numbers and that the semantic fields do not switch across channels. If a label node appears in a path sequence under two different channel numbers and has the same semantic field, the label is marked as a "multi-channel shared semantic label." Duplicate path numbers are assigned to such labels, and a cross-path table is generated. After processing all nodes, a mapping list is created for all path numbers and their corresponding semantic fields. For example, if T011, T012, and T013 are located in P01, P02, and P03, respectively, and their semantic fields are "enter," "execute," and "exit," the path mapping is established as P01→enter, P02→execute, and P03→exit. Finally, the semantic field path relationship chain of all label nodes is output in path number order, resulting in a semantic field path mapping relationship.

[0072] S202: calling the semantic field path mapping relationship, screening the path number sequence that is not continuously corresponding to the semantic field, identifying the label nodes where the semantic correspondence relationship is interrupted, extracting the identification values ​​of the interrupted nodes in the order of the path numbers and classifying and arranging them, to obtain a mapping structure interrupted node sequence;

[0073] Arrange the semantic field records corresponding to each path in ascending order according to the path number, read the semantic field index and its position in the path number sequence path by path, perform a string comparison operation on the semantic field content between the two paths, and use the field value consistency check method. That is, when the semantic fields corresponding to path Pi and path Pi+1 are inconsistent, and the fields do not appear on any other nodes in each other's path sequence, it is determined that there is a semantic interruption between the paths. Continue scanning in the direction of increasing path numbers, and form a number difference sequence for each group of discontinuous path numbers of semantic fields. Path segments with number differences greater than 1 are regarded as structural break segments. The path number difference is recorded as a break span indicator. The allowable range of the break span is set to 1 to 3. If a path segment jumps more than 3 from the previous segment, it is marked as an abnormal break segment, and the label node number corresponding to the break position of the segment is obtained from the path sequence. The label node number format is Txxx, where xxx is a sequential code, for example, T016 represents the 16th label node. The label node is signed and the channel number to which the label node belongs and its corresponding sequential position in the channel are recorded at the same time. The triple format is stored, i.e., (T016, C03, Pos=12), where C03 is the channel number and Pos is the sequential position within the channel. All label nodes that meet the interruption judgment criteria are written into the interruption identification list. During the screening process, it is also necessary to determine whether the label node appears in the semantic field reuse set. If a node appears under two path numbers and its semantic field value remains unchanged, it is not included in the interruption set and is excluded from the interruption identification list. After the screening is completed, all path interruption label numbers are classified and sorted. The mapping relationship between the position of each interruption label node and the corresponding path number is output in ascending order of the path number. The output format is P number → T number sequence. For example, P12 → T034, T037 indicates that two semantic interruption label nodes T034 and T037 appear in the 12th path. Finally, a complete path interruption position index is formed, and a mapping structure interruption node sequence is obtained.

[0074] S203: Based on the mapping structure interruption node sequence, the path numbers and paragraph positions of the included nodes are extracted, the label paths where the interruption points are located are segmented according to the path numbers, and the corresponding path paragraph numbers and interruption node numbers are matched and merged to obtain a semantic interruption structure mapping fragment set;

[0075] Read the path number and node identification value corresponding to the interrupt node, arrange the path number sequence in ascending order, and then use the channel number and path segment number of each interrupt node as a reference to determine the positioning interval of the node in the full path structure. In this interval, identify whether the path number and channel number of the adjacent label nodes before and after it are consistent. If the path numbers of more than three consecutive nodes do not increase or the channel number jumps, the number of the interrupt node in the segment is regarded as a path segmentation point, and a segment index table is generated for all segmentation points. The segment index table records the starting label node number, ending label node number and corresponding path number range of each segment, and then aggregates all label nodes in the segment path where the interrupt node is located. By matching the label node number in the path segment with the interrupt node number, the path segment and interrupt node are merged and determined. The judgment standard is: if the interrupt node number is in the first 30% of the segment path Or the last 30% position, the paragraph is marked as an interruption-prone segment. The proportion division standard is set according to the average length of the entire path. For example: if the path number P21 contains 11 nodes from T011 to T021, of which the interruption node is T013, and T013 is located third in the sequence, accounting for 27.3%, which meets the interruption-prone marking condition, then the paragraph will include all T011 to T021 in the interruption segment aggregation set. On this basis, the nodes marked as interruption-prone segments in all paragraphs are matched again. If the same interruption node number appears between multiple paragraphs, these paragraphs are uniformly classified into the same semantic segment set, and each semantic segment set is sorted twice by the path number. The paragraphs with the same number are merged and the paragraph number label index is updated. Finally, the set of structural fragments aggregated by the path and interruption node number is output to obtain the semantic interruption structure mapping fragment set.

[0076] The specific steps of S3 are:

[0077] S301: Mapping the path break nodes in the segment set based on the semantic break structure, extracting the preceding and following hop connection numbers and path direction sequences corresponding to the nodes, arranging the hop connection structure by path number, and sequentially sorting the hop numbers and comparing them with the path directions to obtain a path hop direction sequence;

[0078] First, read the label number of each broken node and call its hop connection relationship in the original path sequence. The hop connection relationship refers to the jump pair between adjacent label nodes, where one node is the hop starting point and the subsequent node is the end point. After reading the hop connection pair, extract the path number information of each hop. Combined with the order of the previous and next nodes in the path where the hop is located, determine whether the hop connection is a forward jump, that is, whether the node number is monotonically increasing. The hop pairs that do not meet the monotonically increasing condition are marked as reverse hops, and the hop path direction is marked as "reverse order". Then, summarize all hop connection relationships by path number and aggregate them using the path number as the primary key to form a hop set indexed by the path number. In each path hop set, all hop numbers are sorted in ascending order. The order of each hop in the path is compared with its connection direction. If the hop connection order is consistent with the path direction, the consistent direction is recorded as "+1", and if not, it is marked as "-1". Taking path P012 as an example, if its hops are T005→T006 , T006→T007, T007→T006, then the starting point number in the third jump segment is higher than the ending point number, the direction is reversed, and the direction mark is "-1". The direction identification sequence of all jump segments under each path is counted, and a direction identification arrangement vector is formed for each path. Then the direction vectors of all paths are arranged in ascending order according to the path number to form a direction sorting table. The direction sorting table records the overall consistency of the jump segment connection sequence and direction in each path. The direction stability interval is divided according to the number of direction identification changes, and the number of direction changes is divided into Paths with less than one change are marked as direction-continuous segments, and paths with more than two changes are marked as direction-variable segments. The direction stability judgment interval is set as: [0, 1] stable, [2, 3] interrupted, and [4, ∞] unstable. If the jump sequence in path P045 is T010→T011, T011→T010, and T010→T011, and the direction changes twice, it is marked as a direction-interrupted segment. Finally, the jump segment direction identification results under each path are output and uniformly classified to obtain the path jump segment direction sorting.

[0079] S302: Calling the path hop direction sorting function to extract the order of adjacent hops in the path direction sequence, calculating the frequency of repeated occurrence of the path direction in the hop sequence, marking the hop node number combinations with the same repetition frequency, and obtaining the overlap position of the direction sequence;

[0080] The calculation formula for the repetition frequency of the path direction in the hop sequence is:

[0081]

[0082] Among them, F rep represents the frequency of repetition of the path direction in the hop sequence, D z Represents the cumulative number of occurrences of path direction z in the hop sequence, represents the mean number of occurrences of the path direction, G z represents the mean square sum of the differences between the position numbers of adjacent hops in the path in direction z, V z represents the number of path segments corresponding to direction z in the hop sequence, R z The distribution range of the path segments in direction z covers the number of hops, L z represents the number of hop label nodes in the path segment where direction z is located, and u represents the total number of path directions in the hop sequence.

[0083] Assumptions:

[0084] The total number of sample paths is 500. After sequence extraction and jump segment extraction, a total of 5 categories of valid path directions are identified, that is, u = 5;

[0085] The following statistical values ​​were obtained during actual monitoring:

[0086] D1=12, D2=15, D3=10, D4=8, D5=14;

[0087] It represents the average of the cumulative occurrence times of the five path directions. The calculation process is as follows:

[0088]

[0089] The number group in direction 1 is T101→T104→T107, the difference is 3 and 3, and the mean square sum is G1=3 2 +3 2 =18;

[0090] Similarly, we get: G2=20, G3=10, G4=8, G5=16;

[0091] The count is obtained by extracting the path segment number where the hop segment is located:

[0092] V1=4, V2=3, V3=2, V4=2, V5=3

[0093] It is obtained by the total number of hops in this direction in the path segment, which are:

[0094] R1=10, R2=9, R3=7, R4=6, R5=8

[0095] Through node statistics, we can get:

[0096] L1=30, L2=28, L3=20, L4=18, L5=25

[0097] The intermediate values ​​are calculated as follows:

[0098] Item 1:

[0099] Item 2:

[0100] Item 3:

[0101] Item 4:

[0102] Item 5:

[0103] F rep ≈5.66+9.45+7.97+5.64+9.15=37.87;

[0104] The results show that the repeated distribution of pathway directions in the jump sequence has a heavy tendency, and there is a concentrated overlap in the directions with high overall frequency intensity, which provides a quantitative basis for marking the overlapping positions of subsequent direction sequences.

[0105] S303: Based on the overlapped positions of the direction sequence, node pairs with directional continuity characteristics are screened, path numbers and entry node positions are extracted, and the jump segment path numbers are classified and integrated according to the access sequence to obtain a bridge segment entry path node structure diagram;

[0106] First, extract all hop connection pairs from the path hop direction sorting results. Each connection pair is represented as a combination of the start and end point numbers, and its connection direction identifier is recorded. The direction identifier uses the symbol "+1" to indicate the forward direction and "-1" to indicate the reverse direction. After obtaining all hop connection information, group the connection pairs according to the path number, and put hops with the same path number into the same group. Within each group, compare the values ​​of the adjacent hop connection direction identifiers. If the connection direction identifiers of three consecutive hop groups are consistent and the hop numbers are strictly increasing, then the segment is marked as a direction continuation segment. The marking condition is that the end point number of each hop in the hop sequence is equal to the start point number of the next hop minus one. In the example, if the hop sequence under path P05 is T03→T04, T04→T05, and T05→T06, and the direction identifiers are all "+1", then the segment meets the direction continuation feature. Take T03 as the entry node of this direction segment, record its path number P05 and entry node position number T03, and then continue to check the next hop in the path T06→T0 5. If its direction identifier is "-1", it is considered that the continuation segment has terminated. The traversal continues backward to find the next hop segment chain that meets the continuation criteria. All node pairs that meet the continuation judgment criteria are screened, and the starting node number, path number, and direction status of each pair are recorded. The position index of all entry nodes in the path is sorted, and all starting nodes are regarded as potential bridge segment entrances. The path number and direction status corresponding to each entry node are combined and marked, and combined with the subsequent adjacent hop segment number to form a bridge segment hop segment segment structure. The hop segment segment structure uses the path number as the main index. All hop segment path numbers are classified by direction. Hopping segment chains with the same direction are classified into the same hop segment interval. The entry node, end node, and direction sequence combination formed in each hop segment interval are recorded in the number mapping table. The number mapping table uses the path number as the key and the starting hop segment number and direction sequence as the value. Finally, the path number, node starting point, and connection direction relationship corresponding to all bridge segment entrances are marked and summarized to obtain the bridge segment entrance path node structure diagram.

[0107] The specific steps of S4 are:

[0108] S401: Based on the bridge segment connection nodes in the bridge segment entrance path node structure diagram, extract the jump sequence between adjacent nodes, collect the start and end node numbers and jump directions of the path segment, and combine them in the order of path numbers to obtain a path direction sequence combination;

[0109] First, filter all the node sets marked as entrances from the graph, and record the position number of each entrance node in the path segment to which it belongs. Then obtain the jump relationship between each entrance node and its adjacent nodes. By traversing the path number list, establish a jump pair between each pair of adjacent nodes. The structure of the jump pair is "starting point number-end point number", and judge the jump direction. The judgment standard is whether the starting point number is less than the end point number. If so, the jump direction is marked as "forward", otherwise it is "reverse". Record the identification value of the jump direction and pair it with the path number to generate a pairing group with the structure of "path number-jump direction". Then read the arrangement order of the jump direction in each path, and combine the jump segments with the same consecutive jump direction as a direction segment. The first and last nodes of the direction segment are used as the start and end nodes of the segment respectively. Record the starting number, end point number, path number and direction status of each segment in the structured path table. If there are repeated jumps or reverse jumps, the path segment is not counted in the valid direction segment set. A fault tolerance mechanism is set in the process of screening valid segments, that is, one direction change within the path is allowed to be excluded from the elimination criteria. If the number of direction changes is ≥2, the path segment is eliminated and does not participate in the subsequent combination. For example, if the jump segment sequence in path P014 is T06→T07 (forward), T07→T08 (forward), T08→T07 (reverse), the number of direction changes is 1, and it is still counted as a valid segment. The start and end nodes of the segment are T06 and T08. After all valid direction segments are sorted in ascending order of path numbers, adjacent path segments are sequentially compared. If the end and start points of two path segments are adjacent and have the same direction status, they are merged into a group of combined segments. Then, the identification values ​​of each group of combined segments are used to generate a direction sequence mapping index table. The table records the path number, start and end node numbers, direction identifier, and jump sequence number within the segment, and finally the path direction sequence combination is obtained.

[0110] S402: Calling the path direction sequence combination, extracting the direction identifier and structure level value of the path segment, calculating the frequency of occurrence of the direction identifier in the path sequence, screening the path numbers with the same frequency and the same structure level, and obtaining the direction structure frequency combination path;

[0111] The calculation formula for the frequency of occurrence of direction signs in the path sequence is:

[0112]

[0113] in, Represents the frequency of occurrence of the i-th type direction mark in the path sequence, n i represents the total number of occurrences of the i-th type direction identifier in the current path direction sequence, w ij represents the direction weight value of the i-th type direction mark in the j-th occurrence, l ijRepresents the structural level value of the path segment where the i-th type direction identifier appears for the jth time, m ij represents the path occupancy density of the structural unit connected by the path segment where the direction mark of type i appears for the jth time, s ij Represents the number of adjacent direction identifiers connected to the i-th type direction identifier at the j-th path segment position;

[0114] Assumptions:

[0115] First occurrence (j=1): w i1 =1.2, l i1 =3, m i1 =2.5,s i1 =2;

[0116] Second occurrence (j=2): w i2 =1.0, l i2 =2, m i2 =2.0,s i2 =1;

[0117] The third occurrence (j=3): w i3 =0.8, l i3 =4,m i3 =3.0,s i3 =3;

[0118] 1st time:

[0119] 2nd time:

[0120] 3rd time:

[0121] ∑=0.56+0+1.93=2.49;

[0122]

[0123] Total: 2.49 + 1.732 = 4.222;

[0124]

[0125] The results show that the frequency of occurrence of the i-th type of direction mark in the path sequence is 1.407. This value reflects the distribution characteristics of the direction mark in the path sequence and can be used for subsequent path optimization and analysis.

[0126] S403: Based on the direction structure frequency combined path, extract the jump node number sequence corresponding to the classified path segment, align the node number and path number position according to the direction frequency classification order, and obtain the path slip structure connection sequence diagram;

[0127] First, obtain the set of all path segments classified into the same direction segment from the path direction sequence combination, record the path number and the start and end node numbers corresponding to each path segment in the set, then extract the jump node number according to the jump segment connection order in each path segment, construct a complete node number sequence in the jump segment, and classify all node numbers into the corresponding path number according to the path segment they are in to form a number group, then count the jump direction occurrence frequency of each group of node numbers based on the classified path segment, and accumulate the number of occurrences of the direction identifiers "+1" and "-1" respectively. If a certain direction identifier accounts for more than 80% of the path segment, the direction frequency classification value of the path segment is marked as the direction state corresponding to the direction identifier. This rule is a fixed numerical interval judgment and does not allow fuzzy expression. At the same time, the node numbers are screened for direction consistency by the direction frequency classification value, and only the node numbers with the same direction state are retained for the next step of processing. Then, the node number sequence that passes the screening is sorted in ascending order of the path number, and the sorting result is compared with the original path The order in the path segment number sequence is aligned. The alignment standard is: the order of node numbers must be consistent with the order of their corresponding path numbers. If there is a node number jump and a path number duplication, the group of nodes is marked as a "sequence conflict group" and removed from the path direction sequence. It no longer participates in the connection graph construction. For example, if the node numbers of the jump segment in path P011 are T102, T103, and T104, and their direction frequency is "+1", they are classified into the direction consistent group. If the direction frequency in P012 is "-1" and the node numbers are T101, T100, and T099, the sorting should be T099, T100, and T101. After aligning the original path numbers in ascending order without jumping, they are considered valid groups. A connection index table is generated for all valid node groups according to their path number sequence. The relative position of the node number and the path attribution number are indicated in the index table. Finally, a continuous node sequence map is drawn based on the index table and the direction connection status is indicated to obtain the connection sequence diagram of the path slip structure.

[0128] The specific steps of S5 are:

[0129] S501: Based on the label node group in the path slip structure connection sequence diagram, the corresponding channel aggregation segment number is extracted, the semantic field information of the label node in the aggregation segment is collected, the node sequence with path extension relationship between the fields is identified, and the paths are merged according to the structural continuity to obtain the semantic extension structure node;

[0130] First, identify the path segment number to which each label node belongs and its position index in the path. Arrange all path segment numbers in ascending order and generate a corresponding channel aggregation segment number index list. Then read the label node number in each aggregation segment in turn and call the corresponding semantic field information. The semantic field represents the specific semantic state in text encoding form, such as "enter", "hold", "switch", "exit", etc. All semantic field information is archived according to the aggregation segment number. Then, a sequence operation is performed on the archived semantic field group to determine whether the values ​​of consecutive label nodes in the semantic field form a semantic process chain. A semantic process chain is defined as at least three consecutive nodes whose semantic fields form a logical sequence sequence. For example, if T102 is "enter", T103 is "hold", and T104 is "exit", then these three form a valid semantic chain. If T105 is "enter", T106 is "hold", and T107 is still "hold", then the "exit" end condition is not met and it is considered an incomplete chain and is removed. Further path forwarding is performed on all node groups that meet the semantic order relationship. Path extension relationship analysis determines whether the label nodes in the group have continuous numbering in the original path sequence. If the label numbers in the semantically continuous chain are also continuous, it is considered a path continuous sequence. If there is a gap of more than two hops in the numbering, it is marked as a discontinuous segment. Continuous and discontinuous path segments are marked as S (sequential) and B (broken), respectively, and stored in two path cache tables. The next aggregate segment is scanned and the above operation is repeated. After completion, all S-type path segments are merged and a path segment range index is established based on their start and end label node numbers. B-type segments are then associated with their corresponding original path numbers and checked to see if they can be merged into the S-type segment by forming an end-to-end connection with other B-type segments. The connection condition is: if the semantic field of the end node of the two B segments is "keep" or "switch" and the starting node is "enter", then it is determined that there is a path extension relationship and the B-type segment is merged into the previous S-type path segment. After the merging judgment of this rule is applied to all aggregate segments, the continuous semantic path segments in the merged result are finally recorded as the extension result, resulting in a semantic extension structure node.

[0131] S502: Calling the semantic extension structure node to extract the behavior trajectory number, cluster number, and aggregation level value of the associated node in the path segment, analyzing the corresponding structural features between the trajectory number and the cluster number, and classifying and integrating the label nodes with verified number relationships to obtain the behavior trajectory cluster association sequence;

[0132] First, read the number of each extended node in the path segment and its belonging label information, record the correspondence between the path segment number and the label number, and at the same time obtain the behavior trajectory number of each label node in the path segment. The behavior trajectory number represents the time positioning and hop path of the label in the path evolution. Then read the cluster number corresponding to each node, that is, the index classification position of the label in the label cluster, and then extract the aggregation level value corresponding to the label. The aggregation level value represents the hierarchical level in the cluster structure where the label node is located, which is divided into three fixed segmentation ranges: top layer, middle layer, and bottom layer. For example, L1 is the top layer, L2 is the middle layer, and L3 is the bottom layer. The trajectory number, cluster number and level value are combined to form a ternary index group. A one-to-one relationship check is performed on all triplets. The label nodes with one-to-one corresponding trajectory numbers and cluster numbers in the same path segment are marked to determine whether the corresponding method satisfies the cluster continuity. Cluster continuity is defined as the cluster number continuously increasing or decreasing in the sequence. If the label trajectory number is T101, T102, T103, and the cluster number is G02, G02, G03, it is considered as a continuous cluster label sequence. If it is G02, G04, G03, it is considered as a discontinuous cluster sequence and is removed. Further, the trajectory number position and aggregation level value of the label node marked as continuous cluster are recorded. The node groups with the same aggregation level in the same trajectory segment are screened and merged according to the cluster number. The merging rule is: label nodes with continuous trajectory numbers, consistent cluster numbers, and consistent aggregation levels are grouped together. Each group forms an independent number set. The set format is {T101, G02, L2}, {T102, G02, L2}, {T103, G03, L2}, where the first two items are grouped together, and the latter item is not included in the group due to different cluster numbers. Finally, all label node number sequences that meet the cluster number consistency and behavior trajectory continuity are uniformly output to obtain the behavior trajectory cluster association sequence value.

[0133] S503: Based on the behavior trajectory cluster association sequence, extract the label node number group with consistent structural hierarchy in the path segment, combine and associate the behavior trajectory with the node number, regularize the node behavior with association characteristics in the channel aggregation area, and obtain the semantic behavior focus analysis set;

[0134] First, each label node in the sequence is screened, and the path segment number and corresponding aggregation level value of the node are read. The label nodes are grouped with the aggregation level as the primary key, and the nodes with the same aggregation level after grouping are extracted as a structural consistency set. Then, the behavior trajectory number of each label node in the structural consistency set is read. The trajectory number represents the temporal activity identifier of the node in the path segment. Then, the trajectory number and the label number are combined in ascending order, and the path segment number is used as an auxiliary index to form a trajectory-label pair. Position alignment operation is performed on all trajectory-label pairs to determine whether the trajectory number is monotonically increasing. If the number jump does not exceed 2, it is retained. Otherwise, the trajectory pair is eliminated and not included in the subsequent integration. At the same time, it is determined whether the trajectory numbers in each aggregation level group overlap, such as in two If the same trajectory number appears between discontinuous nodes, it is regarded as a behavioral split fragment. Its path number is recorded and set as a conflict identifier. It does not participate in the regularization operation. After the trajectory number cleaning is completed, all trajectory numbers and node numbers are bound to form a behavioral association group, which is classified into a channel aggregation segment according to the path segment number. The behavioral association group in each aggregation segment is used as the basic behavioral unit. The semantic field value and channel number of the node number in each unit are read to determine whether the same trajectory number is distributed in multiple channel numbers. If a cross-channel jump occurs but the semantic field remains consistent, it is marked as "cross-channel semantic continuity". Such nodes are reorganized and sorted according to the channel number and uniformly included in the regularization process. Finally, the behavioral label set regularized by the four elements of path, channel, hierarchy and semantics is output to obtain a semantic behavior focus analysis set.

[0135] See also Figure 2 , a big data information analysis system based on artificial intelligence, including:

[0136] The direction cluster identification module obtains the direction vector of the jump structure node, extracts the channel number and path number of the mutation point label node, combines the direction vector sequence with the channel arrangement position, combines the nodes with continuous directions, aggregates the path numbers according to the jump point order, and obtains the direction jump aggregation label group map;

[0137] The semantic mapping link breaking module aggregates the node paths in the tag group graph based on the direction jump, reads the semantic annotation fields and path segment numbers of the tags in the aggregation segment, merges the path numbers of nodes with discontinuous upstream and downstream semantic fields, and obtains a set of semantic interruption structure mapping fragments;

[0138] The pathway connection and reorganization module maps the broken nodes in the fragment set based on the semantic interruption structure, extracts the jump segment connection number and pathway direction, compares the connection node position and direction sequence, groups the jump segments with structural connection, and obtains the bridge segment entrance path node structure diagram;

[0139] The sliding structure linkage module extracts the jump sequence to form a path direction sequence based on the connection nodes in the bridge section entrance path node structure diagram, calls the path offset trend and structure level value, combines the path segments with consistent structure, constructs a continuous jump path, and obtains the path sliding structure connection sequence diagram;

[0140] The behavioral feature focusing module connects the label nodes in the sequence diagram based on the path slip structure, identifies the semantic extension structure of the channel aggregation segment, calls the behavioral trajectory field, cluster relationship and aggregation level, summarizes the node groups with overlapping clusters and consistent structures, and obtains the semantic behavior focusing analysis set.

[0141] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A big data information analysis method based on artificial intelligence, characterized in that: The following steps are involved: S1: Obtain the jump point direction vector for jump structure analysis in the trajectory, compare the channel number of the direction mutation point with the path number and the order of adjacent jump points, merge the label nodes according to the direction consistency, and obtain the direction jump aggregation label group map; S2: Based on the identified node path in the direction jump aggregation label group graph, according to the label node in the aggregation segment, read the semantic track field and path number, split the path segment without upstream and downstream mapping relationship, and obtain the semantic interrupt structure mapping fragment set; S3: Based on the break nodes in the semantic interruption structure mapping fragment set, the connection numbers and direction sequences of the forward and backward jump segments are analyzed, and the discontinuous segments with consistent directions are marked as bridge segment entrances to obtain a bridge segment entrance path node structure diagram; S4: Based on the bridge segment connection relationship in the bridge segment entrance path node structure diagram, the jump direction sequence is compared, and the path segments with consistent offsets and continuous levels are reordered to obtain a path slip structure connection sequence diagram; S5: Based on the label node group in the path slip structure connected sequence diagram, the semantic extension structure in the channel aggregation area is analyzed, and the continuous trajectory and semantic distribution are merged into behavior segments to obtain a semantic behavior focus analysis set.

2. The big data information analysis method based on artificial intelligence according to claim 1 is characterized in that: The direction jump aggregation label group map includes a jump point direction vector sequence, a channel number sequence mapping set, a direction convergence label node group, and a path structure branch paragraph set; the semantic interruption structure mapping fragment set includes a path paragraph number sequence, a semantic field broken link node set, a discontinuous mapping path index, and a structure interruption node classification result; the bridge section entry path node structure diagram includes a jump section connection number index, a path direction sequence set, a connectable path fragment set, and a bridge section entry node identification set; the path slip structure connection sequence diagram includes a jump direction offset sorting group, a structure hierarchy continuous node group, a path linkage slip segment number sequence, and a channel sequence mapping table; the semantic behavior focus analysis set includes a behavior trajectory aggregation sequence, a cluster number classification table, a path hierarchy relationship set, and a semantic pattern structure item.

3. The big data information analysis method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of S1 are: S101: Obtain a node set of the jump structure in the trajectory, extract the direction vector of the node and the connectivity sequence between adjacent nodes, combine the arrangement trend of the direction vector with the node jump position sequence, analyze the direction structure differences of the label nodes in the direction mutation area, and obtain the jump point direction offset relationship; S102: Based on the jump point direction offset relationship, the channel number and path number of the label node where the mutation point is located are extracted, the channel number arrangement order and the number position of the jump point in the path are combined, and a channel node sequence with consistent direction offset trend is selected and sequentially integrated to obtain a channel label sequence structure; S103: Based on the channel label sequence structure, the path numbers and channel structure sequences to which the label nodes belong are segmented by number, label combinations with similar directions in the same path are collected and the path structure is split, the obtained label node sequences are branched and clustered to obtain a directional jump aggregation label group map.

4. The big data information analysis method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of S2 are: S201: Based on the node path identified in the direction jump aggregation tag group graph, extract the tag node number in the aggregation segment, read the corresponding semantic track annotation field and path segment number, and pair and associate the semantic fields in the order of the path number to obtain a semantic field path mapping relationship; S202: calling the semantic field path mapping relationship, screening the path number sequence that does not correspond to the semantic field continuously, identifying the label nodes where the semantic correspondence relationship is interrupted, extracting the identification values ​​of the interrupted nodes in the order of the path numbers and classifying and arranging them, to obtain a mapping structure interrupted node sequence; S203: Based on the mapping structure interruption node sequence, the path numbers and paragraph positions of the included nodes are extracted, the label path where the interruption point is located is segmented according to the path number, and the corresponding path paragraph numbers and interruption node numbers are matched and merged to obtain a semantic interruption structure mapping fragment set.

5. The big data information analysis method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of S3 are: S301: Based on the path break nodes in the semantic break structure mapping fragment set, extract the preceding and following hop segment connection numbers and path direction sequence corresponding to the nodes, organize the hop segment connection structure according to the path numbers, sort the hop segment numbers in sequence and compare them with the path directions to obtain the path hop segment direction sequence; S302: Calling the path hop direction sorting method, extracting the order of adjacent hops in the path direction sequence, calculating the frequency of repeated occurrence of the path direction in the hop sequence, marking the hop node number combinations with the same repetition frequency, and obtaining the overlap position of the direction sequence; S303: Based on the overlapping positions of the direction sequence, node pairs with direction continuity characteristics are screened, path numbers and entry node positions are extracted, and the jump segment path numbers are classified and integrated according to the access sequence to obtain a bridge segment entry path node structure diagram.

6. The big data information analysis method based on artificial intelligence according to claim 5 is characterized in that: The calculation formula for the repetition frequency of the path direction in the hop sequence is specifically: Among them, F rep represents the frequency of repetition of the path direction in the hop sequence, D z Represents the cumulative number of occurrences of path direction z in the hop sequence, represents the mean number of occurrences of the path direction, G z represents the mean square sum of the differences between the position numbers of adjacent hops in the path in direction z, V z represents the number of path segments corresponding to direction z in the hop sequence, R z The distribution range of the path segments in direction z covers the number of hops, L z represents the number of hop label nodes in the path segment where direction z is located, and u represents the total number of path directions in the hop sequence.

7. The big data information analysis method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of S4 are: S401: Based on the bridge segment connection nodes in the bridge segment entrance path node structure diagram, extract the jump sequence between adjacent nodes, collect the start and end node numbers and jump directions of the path segment, and combine them in the order of path numbers to obtain a path direction sequence combination; S402: Calling the path direction sequence combination, extracting the direction identifier and structure level value of the path segment, calculating the frequency of occurrence of the direction identifier in the path sequence, screening the path numbers with the same frequency and the same structure level, and obtaining the direction structure frequency combination path; S403: Based on the directional structure frequency combined path, extract the jump node number sequence corresponding to the classified path segment, align the node number and path number position according to the directional frequency classification order, and obtain a path slip structure connection sequence diagram.

8. The big data information analysis method based on artificial intelligence according to claim 7 is characterized in that: The calculation formula for the frequency of occurrence of the direction identifier in the path sequence is specifically: in, Represents the frequency of occurrence of the i-th type direction mark in the path sequence, n i represents the total number of occurrences of the i-th type direction identifier in the current path direction sequence, w ij represents the direction weight value of the i-th type direction mark in the j-th occurrence, l ij Represents the structural level value of the path segment where the i-th type direction identifier appears for the jth time, m ij represents the path occupancy density of the structural unit connected by the path segment where the direction mark of type i appears for the jth time, s ij Represents the number of adjacent direction identifiers connected to the i-th type direction identifier at the j-th path segment position.

9. The big data information analysis method based on artificial intelligence according to claim 1 is characterized in that: The specific steps of S5 are: S501: Based on the label node group in the path slip structure connection sequence diagram, extract the corresponding channel aggregation segment number, collect semantic field information of the label node in the aggregation segment, identify the node sequence with path extension relationship between the fields, merge the paths according to structural continuity, and obtain the semantic extension structure node; S502: Calling the semantic extension structure node, extracting the behavior trajectory number, cluster number, and aggregation level value of the associated node in the path segment, analyzing the corresponding structural features between the trajectory number and the cluster number, classifying and integrating the label nodes whose number relationships have been verified, and obtaining the behavior trajectory cluster association sequence value; S503: Based on the behavior trajectory cluster association sequence value, extract the label node number group with consistent structural hierarchy in the path segment, combine and associate the behavior trajectory with the node number, regularize the node behavior with association characteristics in the channel aggregation area, and obtain the semantic behavior focus analysis set.

10. A big data information analysis system based on artificial intelligence, characterized in that: The system is used to implement the artificial intelligence-based big data information analysis method according to any one of claims 1 to 9, and the system includes: The direction cluster identification module obtains the direction vector of the jump structure node, extracts the channel number and path number of the mutation point label node, combines the direction vector sequence with the channel arrangement position, combines the nodes with continuous directions, aggregates the path numbers according to the jump point order, and obtains the direction jump aggregation label group map; The semantic mapping link breaking module jumps the node path in the aggregation tag group graph based on the direction, reads the semantic annotation field and path segment number of the tag in the aggregation segment, merges the path numbers of the nodes with discontinuous upstream and downstream semantic fields, and obtains a semantic interruption structure mapping fragment set; The path connection and reorganization module extracts the jump segment connection number and path direction based on the broken nodes in the semantic interruption structure mapping fragment set, compares the connection node position and direction sequence, groups the jump segments with structural connection, and obtains the bridge segment entrance path node structure diagram; The sliding structure linkage module extracts the jump sequence to form a path direction sequence based on the connection nodes in the bridge section entrance path node structure diagram, calls the path offset trend and the structure level value, combines the path segments with consistent structures, constructs a continuous jump path, and obtains the path sliding structure connection sequence diagram; The behavior feature focusing module connects the label nodes in the sequence diagram based on the path slip structure, identifies the semantic extension structure of the channel aggregation segment, calls the behavior trajectory field, cluster relationship and aggregation level, summarizes the node groups with overlapping clusters and consistent structures, and obtains the semantic behavior focusing analysis set.

Citation Information

Patent Citations

  • Road surface identification processing method, device and equipment

    CN118111412A

  • Bus route generation method and device, equipment and medium

    CN118779390A

  • Urban safety knowledge graph construction method based on AI large model and related device

    CN119623603A

  • Method and system for collecting network security threat information

    CN119743335A

  • Urban traffic mode recognition system based on trajectory clustering

    CN119939325A

Cited By

  • Intelligent data processing method and system based on knowledge graph

    CN121051252A

  • Intelligent checking and circulation method for business expansion application data oriented to multi-role cooperation

    CN121146705A

  • Interactive manual data management method and system

    CN121233692A

  • User behavior analysis and refined operation method based on big data

    CN121235294A