Data integration method and system of consultation service platform based on cloud computing

Through the cloud-based data integration method, the inconsistent data standards and insufficient recommendation accuracy in multi-source heterogeneous data integration of the consulting service platform are solved, and unified data adaptation and personalized recommendation are achieved, which improves the efficiency of data integration and recommendation accuracy.

CN120353898AInactive Publication Date: 2025-07-22TONGHUI TECH TRANSFER (ZAOZHUANG SHANTING DISTRICT) CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510678555.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing consulting service platforms have problems such as inconsistent data standards, difficulty in semantic fusion and insufficient recommendation accuracy in multi-source heterogeneous data integration and recommendation services, especially in the lack of adaptation of voice features among cross-region users and cross-original field conflict analysis.

Method used

The cloud-based data integration method is adopted to achieve unified data adaptation and personalized recommendation through voice and color standardization, dynamic adaptation data fusion, space-time consistency verification, entity relationship reasoning and multi-layer perception personalized recommendation.

Benefits of technology

It improves the efficiency and consistency of data integration, enhances semantic parsability and adaptability of recommended content, and significantly improves the accuracy and user acceptance of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353898A_ABST
    Figure CN120353898A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data integration, in particular to a data integration method and system of a consultation service platform based on cloud computing. The method comprises the following steps: acquiring multi-source heterogeneous data of a consultation service platform, and performing dynamic adaptive data fusion to obtain unified adaptive data; performing space-time consistency verification according to the unified adaptive data to obtain service platform integrated data; performing entity relationship reasoning on the service platform integrated data to obtain an entity relationship reasoning map; performing window-period platform consultation mode identification based on the entity relationship reasoning map to obtain service platform consultation mode data; and performing multi-layer perception personalized recommendation according to the consultation mode data of the service platform to obtain a dynamic content recommendation data set, and uploading the dynamic content recommendation data set to the consultation service platform to execute a content recommendation task. According to the method, the high efficiency of data integration, the semantic integrity of information expression and the availability of recommended content are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data integration, and particularly to a data integration method and system for a consulting service platform based on cloud computing. Background Art

[0002] Existing consulting service platforms usually need to integrate multi-source heterogeneous data, including user historical consulting records, third-party knowledge base data, expert feedback information, industry dynamic data, and real-time session data, etc. The above data sources have problems such as large differences in data structures, diverse data types, and inconsistent update frequencies, which pose great challenges to subsequent unified processing, analysis, and intelligent recommendation. Traditional methods mostly adopt manual configuration of interfaces or data extraction methods based on ETL (Extract-Transform-Load) tools to uniformly import various data into a local database. However, traditional methods lack dialect tone reconstruction and gender timbre feature migration for voice data, resulting in insufficient adaptation of cross-regional user voice features. In addition, the parsing of cross-source field conflicts in traditional methods usually relies on manually preset rules and lacks a dynamic parsing mechanism, resulting in a relatively high redundancy rate of cross-source fields and low data consistency after merging. The misjudgment rate of field value range conflicts remains at a high level for a long time, seriously affecting the integrity of data integration. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a data integration method and system for a consulting service platform based on cloud computing to solve at least one of the above technical problems.

[0004] To achieve the above object, a data integration method for a consulting service platform based on cloud computing includes the following steps: Step S1: Obtain multi-source heterogeneous data of the consulting service platform, and perform voice timbre standardization on the multi-source heterogeneous data of the consulting service platform to obtain voice timbre standardized heterogeneous data; perform dynamic adaptation data fusion on the voice timbre standardized heterogeneous data to obtain unified adaptation data; Step S2: Perform spatio-temporal consistency verification according to the unified adaptation data to obtain integrated data of the service platform; Step S3: Perform entity semantic level division on the integrated data of the service platform to obtain entity level data; perform entity relationship reasoning according to the entity level data to obtain an entity relationship reasoning graph; Step S4: Analyze user consulting preferences based on the entity relationship reasoning graph, and perform window period platform consulting mode recognition according to the user consulting preferences to obtain service platform consulting mode data; Step S5: Perform multi-layer perceptron personalized recommendation based on the service platform consulting mode data to obtain a dynamic content recommendation data set, and upload it to the consulting service platform to execute a content recommendation task.

[0005] In view of the problems existing in the existing consulting service platforms, such as inconsistent data standards, difficult semantic fusion, and insufficient recommendation accuracy in multi-source heterogeneous data integration and recommendation services, the present invention proposes a full-process data fusion and intelligent recommendation processing system. This method improves data traceability and classification clarity by configuring unique identifiers and structural metadata for heterogeneous data sources, avoiding field conflicts and redundant interference. In the processing of voice fields, by combining dialect and gender voice characteristics recognition and migration reconstruction, the standardization of voice content is realized, and the consistency of data structure and semantic parsability are improved. In terms of field fusion, through field type normalization, semantic similarity threshold setting (0.75–0.95), value range sampling (ratio 0.8), and conflict scoring mechanism (interval [0,1]), the fusion quality and efficiency are effectively improved. The spatio-temporal verification solves the problem of aligning heterogeneous data through unified time mapping and spatial scale normalization. In the semantic modeling stage, the word length, word segmentation granularity, and frequency threshold are limited to enhance the accuracy of entity recognition, and a high-confidence entity relationship graph is constructed through context encoding and triple reasoning. In the recommendation part, a behavior tensor and a dynamically weighted behavior perception matrix are constructed, and combined with a three-channel factor model and an online sorting mechanism, the accurate matching of recommended content and user behavior is realized, significantly improving the recommendation adaptability and the intelligent service level of the platform.

[0006] Optionally, the present specification also provides a data integration system for a consulting service platform based on cloud computing, which is used to execute the data integration method for the consulting service platform based on cloud computing as described above. The data integration system for the consulting service platform based on cloud computing includes: A multi-source heterogeneous field fusion module, configured to obtain multi-source heterogeneous data of the consulting service platform, perform voice and sound standardization on the multi-source heterogeneous data of the consulting service platform to obtain voice and sound standardized heterogeneous data; perform dynamic adaptation data fusion on the voice and sound standardized heterogeneous data to obtain unified adaptation data; A spatio-temporal consistency verification module, configured to perform spatio-temporal consistency verification based on the unified adaptation data to obtain integrated data of the service platform; An entity relationship reasoning module, configured to perform entity semantic level division on the integrated data of the service platform to obtain entity level data; perform entity relationship reasoning based on the entity level data to obtain an entity relationship reasoning graph; A consulting mode recognition module, configured to analyze user consulting preferences based on the entity relationship reasoning graph, and perform window period platform consulting mode recognition according to the user consulting preferences to obtain service platform consulting mode data; A personalized recommendation module, configured to perform multi-layer perception personalized recommendation based on the service platform consulting mode data to obtain a dynamic content recommendation data set, and upload it to the consulting service platform to execute the content recommendation task.

[0007] The data integration system of the cloud computing-based consulting service platform of the present invention can implement any data integration method of the cloud computing-based consulting service platform of the present invention, and is used as a medium for coordinating operations and signal transmission between various modules to complete the data integration method of the cloud computing-based consulting service platform. The internal modules of the system cooperate with each other to ensure the personalized adaptation of recommended content and the matching optimization of platform resources, thereby greatly improving the usability and user acceptance of the recommended content. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 It is a schematic flowchart of the steps of the data integration method of the cloud computing-based consulting service platform of the present invention; Figure 2 It is a detailed schematic flowchart of step S1 in the present invention; The implementation, functional features, and advantages of the object of the present invention will be further described with reference to the accompanying drawings in conjunction with embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0009] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0010] To achieve the above object, please refer to Figures 1 to 2 , the present invention provides a data integration method for a cloud computing-based consulting service platform, and the method includes the following steps: Step S1: Obtain multi-source heterogeneous data of the consulting service platform, perform voice and color standardization on the multi-source heterogeneous data of the consulting service platform to obtain voice and color standardized heterogeneous data; perform dynamic adaptation data fusion on the voice and color standardized heterogeneous data to obtain unified adaptation data; In this embodiment, user historical consultation logs (JSON structure), data returned by third-party knowledge interfaces (XML format), industry data subscription streams (Kafka streaming data), and platform expert feedback records (MySQL structured data) are selected as heterogeneous data sources through a consultation service platform. By constructing a structure adaptation mapping module based on source identifiers, a unique source tag and structure meta-information are assigned to each type of data in sequence. For example, JSON fields are parsed into flat key names in a nested path manner. A field extraction rule library is used to perform structure normalization processing on the data stream. A matching template is set for each type of field, and it is uniformly converted into the Parquet columnar format supported by the platform for storage. Then, a field semantic alignment table is constructed, and fields are merged according to the semantic similarity of field names (such as name similarity > 0.75) and field value types (numeric, enumeration, boolean). For example, the fields "industry_tag" and "sector_label" are merged into "industry_code" in multiple data sources. At the same time, a field source trust sorting strategy is executed for field pairs with the frequency of field conflict records ≥ 10 times. Finally, all structure-normalized data are dynamically merged through a field-level fusion strategy to generate unified adapted data, and the data volume compression rate is controlled within 20% - 30%.

[0011] Step S2: Perform spatio-temporal consistency verification based on the unified adapted data to obtain the integrated data of the service platform; In this embodiment, the spatio-temporal consistency check mainly includes two sub-modules: a time window synchronization check module and a spatial location validity identification module. The processing flow and key technical means are as follows: A. The time consistency check module includes: 1) Time field extraction and standardization processing: Extract the "query_time" field from the unified adaptation data and convert it into a standard timestamp (in UTC format, accurate to the second); Use the built-in datetime library and pytz library in Python to perform standard time zone normalization processing on all time data (unify and convert it to Beijing time in the eighth time zone). 2) Time window reorganization and consistency discriminant parameter setting: Set the platform-level time consistency tolerance window to ±30 seconds and adopt a bilateral sliding window mechanism; Use the sliding window function to construct a continuous time period for each user ID (such as the "user_id" field), with a window length of 5 minutes and a sliding step of 60 seconds; Sort the events within the sliding window by time, calculate the time difference between adjacent events, and if the time difference is greater than the set threshold (such as 1800 seconds) or crosses the day boundary (date line), mark it as time anomaly data and exclude it. 3) Time period aggregation and sparse segment cleaning: Group the time data that passes the consistency check by user, and use the aggregation function to merge it into a "time stable segment", recording the start time, end time, and number of events within this segment; Exclude the time periods with the number of events less than 3 (the minimum record number of valid time period data set by the platform). B. The spatial location validity identification module includes: 1) Geographic coordinate standardization and precision control: Extract the field "user_location" (latitude and longitude in string format, such as "116.391,39.907"); Use the geopy library in Python for format parsing and retain the latitude and longitude to 4 decimal places (precision controlled within the range of ±0.0005 degrees, corresponding to an actual spatial resolution of about 50 meters). 2) Reverse geocoding and regional matching mechanism: Call the Amap API (http: / / restapi.amap.com / v3 / geocode / regeo) for all standard latitude and longitude coordinates to reverse resolve the coordinates into district and county-level units (field "district_code"); Use the built-in service area white list to perform a look-up match on the parsed district and county-level unit code field. If it is not within the supported area range, it is determined as spatially inconsistent data and directly excluded.

[0012] Step S3: Perform entity semantic level division on the service platform integrated data to obtain entity level data; Perform entity relationship reasoning based on the entity level data to obtain an entity relationship reasoning graph; In this embodiment, text fields (such as "problem description", "user tags") and structured interaction records (such as click items, feedback content) are extracted from the integrated data of the service platform. A tokenizer constructed based on a context dictionary is used for word segmentation and named entity recognition to identify explicit entities (such as "tax planning", "merger and acquisition process"). The identified entity set is encoded into 768-dimensional context embedding vectors through the RoBERTa pre-trained model. By calculating the cosine similarity, entity pairs with a similarity > 0.85 are clustered and merged to form semantic aggregation clusters. Using the aggregation clusters as nodes, a semantic hierarchy graph with a hierarchical depth of 4 levels is constructed. For example, from "financial management" to "financial statements" and then to "consolidated statements" and "profit distribution". The hierarchical division is completed based on the semantic coverage index (the number of entity aggregations in a single layer needs to be ≥ 8) for hierarchical indexing. From the structured interaction records, the user ID and interaction time are extracted as the aggregation primary keys. With a window of 5 records and a sliding window step of 3 records, the co-occurrence times of entity pairs are counted for each window. Only entity pairs with a clear hierarchical path (parent-child or same-group clustering relationship) are retained, and unassociated entity groups are excluded. For co-occurring entity pairs, the frequency threshold is set to more than 3 times, and the upper and lower position relationships are marked according to the hierarchical relationship direction (for example, the upper-layer clustering number < the lower-layer number). The results are filled into the entity co-occurrence frequency matrix M, where M[i,j] contains fields: entity pair identifier (i,j), co-occurrence times count_ij, relationship type (upper / lower / sibling), and hierarchical difference L_diff. The entity pairs in the co-occurrence matrix are mapped to their RoBERTa vector representations, and the entity pair context vector pairs (A_vec, B_vec) are extracted. After concatenating the vectors, they are input into a three-layer multi-sensory scoring structure (MLP), which outputs three types of scores respectively: S1 = semantic similarity score, S2 = normalized co-occurrence frequency score, S3 = hierarchical consistency score. The final edge weight calculation formula is: wi,j = 0.4×S1 + 0.3×S2 + 0.3×S3. The edge weight retention threshold is set to 0.6, and entity pairs with edge weights greater than the threshold are selected as the potential entity relationship edge set, and the output format is: {entity A, entity B, weight, semantic label pair}. In the entity semantic hierarchy graph, path reachability analysis is performed on all entity pairs in the potential relationship edge set using depth-first traversal (DFS), with the maximum number of hops limited to 5. All intermediate nodes in the path are extracted, and the path ID, number of hops, and node sequence are recorded. The path structure is matched with the existing entity triple relationship templates in the consulting platform knowledge base. The template structure is such as {"enterprise" - "investment" - "project"}, and each type of template comes with a structure pattern and role mapping. The Jaccard structure similarity is used to calculate the structural consistency between the path and the template, and those with a similarity ≥ 0.7 are retained and entered into the candidate path set. For paths that match the template, operations such as entity identification alignment (such as unifying IDs) and role normalization (such as classifying "Company A" as "enterprise") are performed to synthesize the triple reasoning results. The condition for adopting the triple is set that the synthesis effectiveness ≥ 85%.For the candidate triple inference results, confidence calculation is performed. The comprehensive confidence Score is set as follows: Score = 0.5 × template fitness + 0.3 × semantic consistency + 0.2 × hierarchical relationship stability; only triples with a confidence ≥ 0.65 are retained to form the final semantic relationship table. Conflict detection and elimination are performed on triples with conflicts (such as an entity "belonging to" two non-intersecting nodes). Path completion is performed on missing path triples (such as the start and end entities exist in the graph but the intermediate path is missing). The completion method is to find relay entities with a hop count ≤ 3 in the existing co-occurrence frequency matrix and verify the synthesized path structure. If the score ≥ 0.65, it is completed as a legal path. All triples in the final semantic relationship table are mapped to the semantic hierarchy graph, and the edge type (such as "belong to", "contain", "associate", etc.) is set as the directed edge label. The edge attributes include weight, confidence, and path source ID. NetworkX is used to construct a complete inference graph, and topological connectivity check, loop path elimination, and node redundancy merging processing are performed on the graph. Finally, an entity relationship inference graph is generated.

[0013] Step S4: Analyze the user's consultation preferences based on the entity relationship inference graph, and identify the consultation mode of the window period platform according to the user's consultation preferences to obtain service platform consultation mode data; In this embodiment, based on the generated entity relationship graph, taking the user's historical query trajectory as the main line, a node sequence of entities that the user has previously concerned about is extracted (such as "registered capital" - "financing method" - "capital reserve"), and the interaction frequency and time distribution are marked. For each record, according to the semantic level and relationship path to which the consulted entity belongs, it is encoded into a multi-dimensional feature vector, where the entity level is used as the position embedding dimension, the interaction frequency is used as the weight factor, and the interaction timeliness is used as the time decay factor for weighted superposition. Subsequently, the platform takes the high-frequency interactive entities of the user in the most recent consecutive 7 days as the center, and performs weighted averaging on all entity vectors according to a time sliding window to form a dynamic interest core; and combines static information such as the industry type, enterprise scale, and historical consultation preference tags in the user portrait for splicing and fusion, and finally constructs a comprehensive user preference vector model including semantic structure preference, time behavior characteristics, and content sensitivity. On this basis, the observation window is set to 7 consecutive days, focusing on the distribution of active consultation topics of the user during this period, and analyzing its topic jump pattern and the semantic jump path length between entities to obtain the user's consultation preference. The platform introduces high-frequency consultation tags (such as "tax", "financing", "compliance") as pattern recognition anchor points, identifies the path set of multi-entity jumps of the user under the same tag, and determines it as a highly reusable consultation path on the platform. Combining the user portrait (enterprise scale, industry) and consultation intention tags, a typical consultation mode path is constructed (such as medium-sized technology enterprises are concerned about "equity design → option incentive → salary structure"), and the output is the consultation mode data of the platform, with a coverage rate of more than 65% of the historical user consultation path, and the average length of the jump path is not less than 3.

[0014] Step S5: Perform multi-layer perception personalized recommendation based on the consultation mode data of the service platform to obtain a dynamic content recommendation data set, and upload it to the consultation service platform to execute the content recommendation task.

[0015] In this embodiment, the identified consultation mode data is matched with the current user portrait, and the matching threshold is set to a vector angle < 30°, and a context-aware template is dynamically constructed according to the user's current access intention (such as clicking on keywords, search behavior). According to the template, the recommended level is selected as 3 layers, corresponding to: approximate semantic content recommendation, high-scoring content recommendation for similar users, and popular content recommendation for platform experts. The number of recommended contents in each layer is not less than 5, and the recommended contents need to meet the content update cycle of no more than 30 days and the user score of greater than 3.8. The content data is extracted from the platform content pool, and a source credibility label and a timeliness score are attached. The final recommendation result is submitted to the platform recommendation engine in a structured JSON format, pushed to the user interface in real time through WebSocket, and each recommended click feedback is recorded for model update. The push delay of the recommended content needs to be controlled within 1.5 seconds, and it supports concurrent pushing of more than 1000 user visits per second.

[0016] Optionally, step S1 is specifically as follows: Step S11: Obtain multi-source heterogeneous data of the consulting service platform, assign a unique source identifier and structure description metadata to each type of data source, and generate an initial heterogeneous data set. In this embodiment, five types of heterogeneous data sources are accessed through the data access gateway module of the consulting service platform, namely: structured user consultation logs (MySQL format), semi-structured expert feedback documents (JSON format), unstructured industry dynamic texts (TXT format), real-time session streams (WebSocket messages), and external knowledge base data (RDF format). A unique source identifier (such as SRC_USER_01, SRC_KB_02, etc.) is dynamically generated for each type of data source, and information such as field structure, data type, and hierarchical nesting relationship is synchronously extracted to automatically generate structure description metadata, which is stored in YAML format. In the access stage, the platform uses the built-in data acquisition scheduler to set the pulling frequency to once every 10 minutes and limit the amount of data collected each time to no more than 5MB to ensure access efficiency and system stability, and finally summarize to form an initial heterogeneous data set.

[0017] Step S12: Perform voice and color standardization on the initial heterogeneous data set to obtain voice and color standardized heterogeneous data, and perform unified field extraction and format normalization processing to generate a structure standardized data set. In this embodiment, all data entries containing voice fields in the initial heterogeneous dataset are extracted, the fields storing audio content are screened out, and the corresponding audio segments are extracted. To ensure the consistency of the input format, the sampling rate of all audio segments is uniformly set to 16 kHz, the number of channels is set to mono, and the resampling function is used for sampling rate conversion. At the same time, the redundant channel data in non-mono audio is removed to generate a standard input audio dataset with a consistent format. Subsequently, a voice activity detection mechanism based on an energy threshold is applied. The mute judgment threshold is set to -40 dB, and the lower limit for continuous voice judgment is set to 300 ms. Background noise and invalid mute segments are removed, and only clear and continuous voice segments are retained to construct a set of effective voice segments for subsequent analysis. For the above set of effective voice segments, frame division is performed in a manner with a fixed frame length of 25 ms and a frame shift of 10 ms, and 20-dimensional Mel-frequency cepstral coefficients (MFCCs) are calculated. At the same time, acoustic feature parameter sets are constructed by combining timbre features such as fundamental frequency (F0), fundamental frequency variation rate (ΔF0), and formant spacing (F1–F2). To distinguish regional pronunciation features in the voice, a pre-constructed regional dialect acoustic feature library is loaded locally, where each type of dialect is characterized by a statistical distribution vector in the form of Di={μMFCC,σMFCC,μF0,σF0}. The frame-level recognition window is set to 500 ms, and the sliding step size is set to 250 ms. The acoustic features are divided according to the window, and the Bayesian similarity distance is used to match with each dialect library. The closest dialect label is extracted to form a distribution confidence vector, and the region corresponding to the maximum confidence value is selected as the dominant pronunciation region, and the dialect recognition label is output. In the gender discrimination process, an empirical division method is used to analyze gender based on the acoustic feature parameter set. The male discrimination interval for the fundamental frequency is [85 Hz–165 Hz], and the female interval is [165 Hz–255 Hz]; the male discrimination interval for ΔF0 is [5 Hz–15 Hz] with a standard deviation less than 8 Hz, and the female interval is [15 Hz–35 Hz] with a standard deviation greater than 10 Hz; the male interval for the formant spacing F1–F2 is [700 Hz–1000 Hz], and the female is [1000 Hz–1400 Hz]. For each voice segment, if the three groups of features simultaneously fall within the interval range of a certain gender, the gender to which the segment belongs is determined and the corresponding gender recognition label is assigned. In the voice style correction process, according to the dialect recognition label, the prosodic features of the original audio are adjusted by combining the tone contour reconstruction function unique to the dialect, including syllable delay, stress elevation, and intonation conversion, etc., to generate an audio reconstruction segment with a unified pronunciation region; according to the gender recognition label, the parameter range of the vocal cord model is adjusted (such as the center of the vocal range frequency band is offset by 5~15 Hz) for vocal range reshaping to generate a vocal range reconstruction segment. Finally, the above two segments are respectively bound to the original data entries, replacing the original audio field content, and output as voice color standardized heterogeneous data.During the process of structure parsing of the initial heterogeneous data set, a field hierarchical parsing module is adopted to expand the nested structure of structured and semi-structured data through field path mapping. For example, "feedback.details.comment" is disassembled into a multi-level field path. In the field extraction link, the platform establishes a unified field dictionary and performs field name mapping and field type matching with reference to the general consultation field set (such as "question_id", "entity_name", "interaction_type", etc.). All numerical fields are uniformly converted to the Float32 format, and date fields are converted to the standard ISO 8601 format. For text fields, UTF-8 encoding standardization is performed, and the character length truncation is set within 512 characters. Finally, all fields are stored in the Parquet format to form a structure-standardized data set.

[0018] Step S13: Construct a cross-source data mapping matrix based on the structure-standardized data set, identify redundant fields and similar semantic units, and perform field merging and priority parsing to generate a semantic mapping data set; In this embodiment, after the construction of the structure standardization data set is completed, first, through the field metadata extraction module, each source data field is uniformly extracted and formatted, and the field name (field_name), data type (data_type), field source identifier (source_id), field value distribution information (value_distribution), and field frequency index (field_freq) are extracted to construct a set of standardized field metadata structure bodies. Subsequently, three types of field similarity measurement operations are performed based on this structure body: First, calculate the field name similarity based on the Levenshtein edit distance, normalize the result to the [0,1] interval, and set the similarity threshold to 0.85; Second, use the overlapping rate of the five - percentile intervals to measure the value range distribution similarity of numerical fields, and a similarity of the value range is considered when the overlapping rate is greater than 70%; Third, for text fields, use the TF - IDF weight vector combined with the cosine similarity, and set the semantic threshold to 0.75. The above three measurement results are weighted and summarized according to the weights of 0.4 (name), 0.4 (value range), and 0.2 (semantics) to form a field similarity matrix, and the field pairs with weighted scores exceeding 0.7 are selected to construct a mapping candidate set. On this basis, an undirected graph structure is used to construct a field association relationship graph, with fields as graph nodes and similar field pairs as edges for connection, and a connected sub - graph decomposition strategy is adopted to identify potential field equivalence classes. For the fields within each equivalence class, according to the field source identifier and frequency index, the field with the highest frequency is selected as the main field, and the rest are marked as candidate fields, and priority labels (P1, P2, P3, etc.) are set in descending order of frequency. Finally, each set of field mapping relationships is stored in the cross - source data mapping matrix in the form of a (main field, candidate field, priority) triple, and is persistently saved in the JSON Lines format to form a semantic mapping data set.

[0019] Step S14: Combine the source identifier and the field priority assignment rule in the semantic mapping data set, perform per - field - level data fusion processing, and perform flow control to generate a fusion - processed data set; In this embodiment, the field fusion scheduling module is used to call the semantic mapping data set, and the following fusion process is performed on each field mapping triple: 1) Streaming window configuration: Set the sliding window parameter to 30 seconds and the sliding step to 15 seconds. The system schedules the fusion processing task every 30 seconds, and the upper limit of the number of field mapping pairs processed in each batch is 20 to control the computing load; 2) Primary and secondary source selection mechanism: Determine the source identifier of the primary field for the current fusion operation (e.g., SRC_USER_01) according to the priority order of the primary field and candidate fields defined in the mapping matrix; 3) Fusion rule setting: a. If the field type is numeric, perform the following judgment: If the difference between the primary and secondary source values is less than the set error threshold (ε = 5% of the primary source mean), take the primary source value; if the difference is greater than the threshold, perform weighted fusion (primary source 0.7, secondary source 0.3); b. If the field is text, perform sentence-level similarity matching (Jaccard similarity threshold 0.6). If the difference is large, adopt the splicing strategy and mark the merged source; c. If the field is time-type data, retain the field with a higher time update frequency as the primary field output; 4) Conflict resolution: When there is a timestamp conflict between the primary and secondary source data, use the data with the latest timestamp as the standard, and write the conflict log into the fusion processing log table. The fusion processing result is cached in the platform data cache area in a row group structure, and the single buffer block does not exceed 64MB. To verify the data integrity, calculate the CRC-32 checksum before and after generating each batch of fusion blocks. After passing the verification, write it into the fusion processing data set, and the storage format is Parquet. During the fusion process, retain the original source information and fusion method record of the field in the following format: {"field_name":"question_subject","primary_source":"SRC_USER_01","secondary_source":"SRC_KB_02","fusion_type":"weighted_merge","merge_confidence":0.89}; Finally, the fusion processing data set can be used as the direct input for generating subsequent unified adaptation data.

[0020] Step S15: Perform outlier removal and keyword field integrity verification on the fusion processing data set to generate unified adaptation data.

[0021] In this embodiment, when cleaning and validating the fused processing dataset, anomaly detection rules are first set for numerical fields, and the upper / lower bounds are set as the historical mean ± 3 times the standard deviation to detect and remove out-of-limit data. Regular validation is performed on string fields. For example, for email format fields, the expression [a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+ is used in this embodiment. Those that do not match are judged as format anomalies and discarded. For key fields (such as "question_id" and "entity_id"), the non-empty rate is required to be ≥ 98%. If the field missing rate exceeds the threshold, the repair process is triggered, and data is backfilled in the mapping matrix using the field association path. After all validations pass, uniformly adapted data with a unified structure, reliable fields, and complete content is finally output as the basic input for subsequent module processing.

[0022] Optionally, the voice sound standardization in step S12 is specifically as follows: Extract the audio segments containing voice content from the initial heterogeneous dataset, set the sampling rate to 16 kHz and mono as the standard input format, and perform resampling and channel unification processing to generate a standardized input audio dataset; In this embodiment, for the content containing voice fields in the initial heterogeneous data, records with audio information are first screened by field name keyword matching (such as "audio", "voice", "media_path"), and the ffmpeg tool is uniformly called to perform audio resampling operation with the sampling rate set to 16 kHz. To maintain data consistency, all audio channels are uniformly processed into mono form, that is, the main channel data is retained for stereo audio. All output files are converted to the WAV format, and the header encoding is set to 16-bit linear PCM encoding to ensure the readability of subsequent processing modules. The standardized input audio dataset generated in this way can be uniquely confirmed through MD5 verification to ensure version control.

[0023] Perform voice activity detection on the standardized input audio dataset, set the silence threshold to -40 dB and the minimum voice duration to 300 ms, and remove background noise and silent segments to generate a set of effective voice segments; In this embodiment, voice activity recognition is performed on the basis of the standardized audio, and the frame-by-frame energy detection method is used to evaluate the silent segments and effective voice segments in each audio segment. During this process, the silence determination threshold is set to -40 dB, and the minimum voice duration is set to 300 ms. The audio is divided into frames with a length of 20 ms and a frame shift of 10 ms, and the duration of consecutive high-energy frames is counted. When the duration requirement is met and the average energy is higher than the set threshold, it is retained as an effective voice segment. Silent segments and high-background-noise segments are assisted in removal by an energy median filter, and finally a set of effective voice segments is formed, and the output is a sub-segment set file with time-scale alignment.

[0024] Set the dimension of Mel Frequency Cepstral Coefficients to 20, frame length to 25 ms, and frame shift to 10 ms. Extract acoustic features from the set of valid speech segments, and simultaneously extract timbre features including fundamental frequency, fundamental frequency variation rate, and formant spacing features, thereby obtaining an acoustic feature parameter set; In this embodiment, for the extracted valid speech segments, acoustic feature extraction is performed. First, a Mel filter bank is used to extract 20-dimensional cepstral parameters as the main spectral features. The frame parameters are set to a frame length of 25 ms and a frame shift of 10 ms. The features extracted include: 20-dimensional MFCC, 1-dimensional fundamental frequency (F0, calculated using autocorrelation period detection), 1-dimensional fundamental frequency variation rate (inter-frame F0 difference), and the spacing between formant frequencies F1 and F2, as timbre discrimination features. Combine the above features to form a frame-level feature matrix in the form of , where N is the number of frames. The matrix file is saved in a standard JSON structure or NumPy format for easy calling in downstream processing.

[0025] Obtain a regional dialect acoustic feature library, and train a dialect recognition model based on the regional dialect acoustic feature library; In this embodiment, import the regional dialect acoustic feature library, which is trained through various dialect corpora. Its core consists of several feature templates, and the template for each dialect contains a mean vector and a covariance matrix. During the training process, aggregate the acoustic parameter sets by dialect category and generate a Gaussian distribution template in the form of , with a dimension of 23.

[0026] Set the frame-level recognition window to 500 ms and the sliding step to 250 ms. Use the dialect recognition model to perform regional speech feature recognition on the acoustic feature parameter set and generate dialect recognition labels; In this embodiment, during the dialect recognition process, set the frame-level window to 500 ms and the sliding step to 250 ms. Calculate the Mahalanobis distance between the feature segments extracted from each time window and all Ti in turn, and assign a regional label based on the minimum distance result. The dominant pronunciation region of the overall audio segment is determined by majority voting of the window labels, and dialect recognition labels are output.

[0027] Obtain a human gender timbre feature library, and perform empirical threshold division according to the human gender timbre feature library. Among them, the male discrimination threshold interval for fundamental frequency is [85 Hz - 165 Hz], and the female discrimination threshold interval for fundamental frequency is [165 Hz - 255 Hz]; The male discrimination threshold interval for fundamental frequency variation rate is [5 Hz - 15 Hz], and the standard deviation of fundamental frequency variation rate < 8 Hz; the female discrimination threshold interval for fundamental frequency variation rate is [15 Hz - 35 Hz], and the standard deviation of fundamental frequency variation rate > 10 Hz; The male discrimination threshold range of formant spacing is [700Hz–1000Hz]; the female discrimination threshold range of formant spacing is [1000Hz–1400Hz]; In this embodiment, a human gender timbre feature library is obtained by referring to relevant materials or expert experience, and statistics are carried out according to the feature differences between males and females in the human gender timbre feature library to divide the gender discrimination threshold ranges of fundamental frequency, fundamental frequency change rate, and formant spacing.

[0028] Based on the empirical threshold, the timbre features in the acoustic feature parameter set are discriminated. When the fundamental frequency, fundamental frequency change rate, and formant spacing simultaneously meet the male / female gender discrimination threshold range, the gender category corresponding to the current speech segment is determined, and a gender recognition label is generated; In this embodiment, it depends on the defined gender discrimination threshold range (empirical threshold). The discrimination criteria include: the fundamental frequency is 85–165Hz for males and 165–255Hz for females; ΔF0 is 5–15Hz and the standard deviation is less than 8Hz for males, 15–35Hz and the standard deviation is greater than 10Hz for females; the F1–F2 spacing is 700–1000Hz for males and 1000–1400Hz for females. Only when all three indicators simultaneously meet a certain gender discrimination interval can a clear mark be made. Samples that cannot be fully matched are marked as "uncertain" and selected for rejection. The recognition label is recorded in the form of One-Hot encoding and written into the data structure in the form of an additional field.

[0029] Perform speech style transfer processing on the acoustic feature parameter set according to the dialect recognition label and gender recognition label, including audio reconstruction of the tone contour in the acoustic feature parameter set according to the dialect recognition label to generate an audio reconstructed speech segment; perform pitch range parameter reconstruction on the timbre features in the acoustic feature parameter set based on the gender recognition label to generate a pitch range reconstructed speech segment; In this embodiment, according to the recognition label, the tone curve and timbre parameters of the audio are reconstructed respectively. In tone reconstruction, the current fundamental frequency trajectory is aligned to a preset tone pattern template (such as the three-tone pattern of Mandarin is 55, 33, 21, and the nine-tone pattern of Cantonese corresponds to the F0 curve sample) using a dialect mapping matrix. The alignment method uses dynamic duration normalization and fundamental frequency smooth projection to achieve tone pattern transformation. In terms of timbre reconstruction, the conversion direction is selected through the gender recognition label, and the pitch range mapping function is called to adjust the F0 center frequency band. For example, the male domain of 135Hz±15Hz is redirected to the female domain of 215Hz±20Hz, and the frequency band mapping ratio is set to β = 0.25 during the operation to control the smoothness of the transformation. Finally, two new audio segments are formed, namely the tone pattern reconstruction result and the pitch range reconstruction result.

[0030] Bind the audio reconstructed speech segments and the pitch reconstructed speech segments to standardize the input audio dataset and semantic labels, and replace the original speech field to generate speech sound and color standardized heterogeneous data.

[0031] In this embodiment, the processed reconstructed speech segments are filled back into the original data structure to replace the original audio field. The semantic label field remains unchanged to ensure that the data semantic association is not affected. The reconstructed data is uniformly saved as a tagged WAV file and a structured JSON description file to form a speech sound and color standardized heterogeneous dataset.

[0032] Optionally, step S13 is specifically as follows: Step S131: Perform a field-level index scan on the structure standardized dataset, extract the field name, type, value range, and occurrence frequency information of each data source, so as to construct a field attribute feature matrix; In this embodiment, based on the structure standardized dataset, first configure the field index scan rule, and set the selection condition as the number of fields ≥ 50 and the data record coverage rate ≥ 85%. For the data sources that meet the conditions, relying on the field parsing module, extract all the structural metadata of the fields based on UTF-8 encoding, including the field name (field_name), field type (limited to four types: string, integer, float, boolean, and fields of other types are automatically discarded), field value range (calculating the upper and lower limits based on quantile(0.025) to quantile(0.975) of the pandas library), and field frequency (the proportion of non-empty records of the field, and low-frequency fields are filtered using a threshold of 1%). All the extracted information is used to construct an attribute feature matrix with rows for each field. The matrix column dimension includes five field dimensions and is stored in the form of a scipy sparse matrix structure to avoid occupying space by low-frequency fields. The field attribute feature matrix is saved in the.npz compressed format for subsequent calls by the semantic similarity analysis module.

[0033] Step S132: Aggregate the semantic and structural features between fields based on the field attribute feature matrix, construct a field semantic similarity matrix, and set a semantic fusion threshold to filter out a set of candidate field pairs; In this embodiment, after loading the field attribute feature matrix, pairwise comparison of fields is performed. The process of constructing the semantic similarity matrix includes three types of feature aggregation: 1) Type consistency comparison: If the field types are exactly the same, it is marked as 1, otherwise it is 0; 2) Value range overlap calculation: For float and integer fields, calculate the ratio of the intersection length to the union length of the normalized upper and lower bounds as the overlap rate, with a range of [0,1]; 3) Semantic label similarity: Perform word segmentation on the field names (using the jieba tokenizer), extract semantic units such as the first word and root words, and calculate the cosine similarity after encoding through a word vector model (word2vec). The three indicators are weighted and summed according to the weights 0.3 (type), 0.4 (value range), and 0.3 (semantic label) to form a comprehensive semantic similarity matrix. Set the fusion threshold to 0.72, only retain the field pairs with scores greater than this value, write them into the candidate field pair set, and save them as a.csv structured file, which includes the field pair ID, similarity value, and the original features participating in the scoring.

[0034] Step S133: According to the candidate field pair set, establish a cross-source field matching mapping table, and combine the source identifier to construct a two-way mapping graph of fields, thereby generating a cross-source data mapping matrix; In this embodiment, read the candidate field pair set, and combine the field source identifier (source_id) provided in the structure-standardized data to establish a two-way mapping path for each field pair, construct a field matching table (field A ⇄ field B), and append semantic explanations (constructed from high-frequency co-occurring words, such as "user_id" ⇄ "uid"). Use the NetworkX graph processing library to construct an undirected graph model, with fields as nodes, matching relationships as edges, and edge weights as similarity scores. For graph structures with more than 200 nodes, adopt a graph clustering and partitioning strategy based on the Louvain method for grouping, and persist the graph structure in the.graphml format. Additional information such as field pair relationships, source mappings, and update period ratios is synchronously written into the cross-source mapping matrix and encapsulated for storage in the JSON structure, with fields: {"primary_field":"user_id","candidate_field":"uid","priority":"P1","source_relation":"A⇄B","update_freq_ratio":1.25}.

[0035] Step S134: Perform field conflict analysis on the cross-source data mapping matrix, and perform field priority parsing and merging according to the field conflict analysis results to obtain a field fusion candidate table; In this embodiment, based on the generated cross-source data mapping matrix, conflict indicators are calculated for each group of fields. 1) Conflict frequency: The difference in the proportion of duplicate values of records in different source fields before and after the merger of the target table is counted. If it exceeds 20%, it is regarded as a high conflict; 2) Value distribution deviation: The Kolmogorov–Smirnov test (KS test) is used to compare the difference in field value distributions. If the D value is greater than 0.3, it is regarded as a serious distribution deviation; 3) Aging overlap coefficient: The field update timestamps are extracted, and the record coincidence rate within the past month is counted. If it is less than 30%, it indicates a large difference in data versions. The three indicators are weighted and comprehensively scored according to the weights of 0.4 (conflict frequency), 0.3 (KS deviation), and 0.3 (aging overlap). If the score is greater than 0.65, it is regarded as a high-conflict field pair, and the lowest fusion priority is given. In the fusion candidate field table, only the field pairs with a score lower than 0.65 are retained, and the main field attribution is determined based on the field source, record completeness rate, and occurrence frequency, and the candidate field triple and score details are output.

[0036] Step S135: Perform consistency verification on the field fusion candidate table, remove the field pairs with a merger conflict rate higher than the set threshold, and perform field-level merger on the remaining field pairs in the field fusion candidate table to generate a semantic mapping data set.

[0037] In this embodiment, after reading the generated field fusion candidate table, consistency checks are sequentially performed on the field pairs, specifically including: 1) Type matching check: Using a strong matching strategy, it is required that the data types of the main field and the candidate field are exactly the same (for example: integer cannot be mixed with float); 2) Value range intersection verification: For numerical fields, the distribution interval is extracted to calculate the intersection ratio. If the intersection ratio is less than 0.4, it is marked as a conflict pair and removed; 3) Conflict rate confirmation: The proportion of data conflict records generated after the field merger is counted. If it exceeds 60% (counting the newly added null values or overwritten values after the merger), then this field pair does not enter the final merger set. The remaining field pairs will perform main field retention, candidate field data migration, and mapping path update through the field merger module. Each merger operation records its source field, merger timestamp, and field retention strategy (such as "main field retention, candidate field migration"), and finally outputs a semantic mapping data set, which is saved in a structured object manner (such as Parquet or JSON Lines) to support subsequent operations such as field fusion, data aggregation, and standard table generation.

[0038] Optionally, step S134 is specifically: Set the sampling ratio of field value range extraction to 0.8, perform value range intersection analysis and time dimension consistency detection on each field pair in the cross-source data mapping matrix, count the conflict frequency, value distribution deviation, and aging overlap coefficient, and generate a field conflict evaluation index set; In this embodiment, a field-by-field pair processing is performed on the cross-source data mapping matrix, and the sampling ratio of the field value range is set to 80%. The specific steps are as follows: 1) Sample extraction: For each field pair, 80% of the data records are randomly extracted from the cross-source data mapping matrix to ensure that the main value distributions of the fields are covered in the samples. 2) Value range intersection analysis: The set intersection calculation is respectively performed on the value sets of the two fields, and the number of intersection elements and their proportions in the two sets are counted. The intersection proportion threshold is set to 30% to mark the field pairs with low coincidence. 3) Time consistency detection: a. Read the record timestamps of the two fields, and slice the field data according to a fixed time window (such as 24 hours); b. Calculate the proportion of the overlapping interval length of the time window. If the overlapping ratio is less than 20%, it is marked as a field with low timeliness consistency. 4) Conflict index calculation: a. Conflict frequency = number of different values / total number of samples, where different values refer to the records with inconsistent values under the same primary key for the field pair; b. Value distribution deviation = |skew1 - skew2|, where skew represents the skewness coefficient of the field value; c. Timeliness overlap coefficient = length of time window overlap / length of time union. The above indexes are all saved in the structure FieldConflictMetrics in the form of floating-point values. The fields include: intersection_rate, conflict_ratio, distribution_skew_delta, time_overlap_rate. Finally, a list of conflict evaluation index sets for all field pairs is output.

[0039] Construct a field conflict feature vector based on the field conflict evaluation index set; In this embodiment, a standardization method is used to perform feature vectorization processing on the field conflict evaluation indexes. The detailed process is as follows: 1) Normalization processing: a. The conflict frequency, skewness difference, and timeliness overlap ratio are uniformly mapped to the [0, 1] interval; b. The normalization values of each type of index are calculated respectively using the Min-Max Scaling method. 2) Weight configuration: a. Set the feature weights: conflict frequency 0.4, distribution deviation 0.35, time overlap 0.25; b. Use the weighted synthesis method to construct a three-dimensional conflict vector [C, D, T] for the field pair, where C, D, and T respectively represent the three normalized indexes. 3) Semantic conflict correction: If the semantic labels of the two fields are inconsistent (such as different units but the same name), perform negative correction by looking up the semantic label comparison table, and the correction amplitude is set to reduce the overall score by 0.1. The constructed feature vector is saved in the form of a JSON array. The fields are: field_a, field_b, conflict_vector (a three-dimensional floating-point array).

[0040] Classify and mark the field conflict patterns using the field conflict feature vector, and set the conflict severity scoring interval to [0, 1] to generate a field conflict level annotation table; In this embodiment, the field pairs are divided into three categories according to the weighted scores of the feature vectors: 1) Calculation of the conflict severity score: Use the formula: Score = 0.4×C + 0.35×D + 0.25×T, and the result is reserved to three decimal places; 2) Classification rules: a) Score ≥ 0.8 → High-risk conflict segment; b) 0.5 ≤ Score < 0.8 → Medium conflict segment; c) Score < 0.5 → Low-risk segment; 3) Annotation generation: Construct a field conflict level annotation table, and the table structure includes: field_a, field_b, conflict_score, conflict_level; all field pairs are stored in the SQLite local database table conflict_levels for query and audit.

[0041] Combine the field occurrence frequencies and data update cycles of field pairs with different conflict levels in the field conflict level annotation table to calculate the source trust weights of the fields, thereby generating a field source weight matrix; In this embodiment, first, count the number of occurrences of each field in its source data source and record it as frequency; calculate the shortest update time interval of the field value in hours and record it as update_interval; then perform the calculation of the source trust weight score of the field: If the field update cycle is less than 24 hours and the occurrence frequency ranks in the top 30%, assign a weight = 1; otherwise, for the remaining fields, use the formula: Weight = min(1, 0.3 + 0.7×(frequency_norm)×(1 - update_interval / 72)) for normalized scoring; The matrix is constructed in the form: field pairs are rows; the data sources corresponding to the fields are columns; the source trust weights corresponding to the fields are values, and the field source weight matrix is output in CSV format.

[0042] According to the field source weight matrix and the field conflict level annotation table, perform priority sorting and merger strategy analysis on each field pair to form a field fusion priority schedule; In this embodiment, the priority sorting process includes: 1) The initial sorting is based on the conflict level: low conflict level first; 2) Calculation of the fusion weight difference factor: If the source weight of a certain field in the field pair is more than 0.2 greater than the other party, then improve its fusion priority; 3) Determination of the fusion strategy type: a) Value range overlap > 50%, update cycle difference < 6 hours → Value merge; b) Semantic labels are the same and the units are the same → Overwrite merge; c) Different names but belonging to the same semantic cluster → Mapping merge; 4) Output the fields of the field fusion priority schedule: field pair, priority number, recommended strategy type, strategy parameters, and the field fusion priority schedule is output in JSONLines format.

[0043] Apply the merging instructions in the field fusion priority schedule to the cross-source data mapping matrix, and mark the merging policy type and confidence label for each field pair, thereby generating a field fusion candidate table.

[0044] In this embodiment, the following fusion processing is performed: 1) Merging scheduling: Execute item by item according to the priority number, first low-conflict, and then high-trust sources; 2) Policy application: a) Overwrite: Retain the field with high weight, and record merge_type=overwrite; b) Value union: Perform the set(field1∪field2) operation, and record merge_type=value_union; c) Semantic alignment: Check the semantic mapping field dictionary, construct a unified field name, and record merge_type=semantic_align; 3) Confidence label: Add a field-level trust mark according to the average weight: high (>0.8), medium (0.5–0.8), low (<0.5); 4) Output format: Each record contains the field pair, the merging policy, the merged field name, and the confidence level, and is stored in the structured field fusion candidate table (fields: field_a, field_b, merged_field, strategy, confidence_label).

[0045] Of particular importance is that the priority sorting and merging policy analysis are specifically as follows: Classify the conflict levels of each field pair in the field conflict level annotation table, perform multi-dimensional scoring according to the conflict frequency, value distribution deviation, and time overlap degree of each field pair, and construct a conflict severity level vector; In this embodiment, first, all field pairs in the field conflict level annotation table are traversed for data rows, and a three-dimensional conflict feature index vector in the form of [conflict_rate, distribution_delta, time_overlap] is constructed for each pair of fields. The specific description is as follows: 1) Conflict frequency (conflict_rate): Count the number of times different values appear for a field pair in the same primary key record, and divide it by the total number of comparable records to obtain the conflict ratio, with the value range being [0, 1]. 2) Value distribution deviation (distribution_delta): Calculate the skewness and kurtosis of the numerical data for the two fields respectively, and then take the absolute difference between the two as the deviation value, which is normalized to [0, 1] using the normalization method. 3) Time overlap degree (time_overlap): Extract the field record timestamps, count the number of records where the values of the two fields appear simultaneously within a 24-hour time window, and divide it by the total number of records within the combined time coverage interval. The result is normalized to [0, 1]. Finally, the three indicators are concatenated into a conflict severity level vector. Each field pair corresponds to a structured record (in JSON format), saving the field pair ID, the scores of the three dimensions, and the vector generation time.

[0046] An initial field priority list is established based on the conflict severity level vector; In this embodiment, the field pairs are comprehensively sorted based on the generated conflict severity level vector. The specific approach is as follows: 1) Weighted scoring formula: Set the weights of each dimension to 0.5 for conflict frequency, 0.3 for value distribution deviation, and 0.2 for time overlap degree; use the formula Score = 0.5×C + 0.3×D + 0.2×T to calculate the comprehensive conflict score for each field pair. 2) Sorting to generate a list: Sort all field pairs in descending order of the comprehensive score to generate an initial field priority list. The list fields include: field pair number, the original scores of the three conflict dimensions, the total score, and the sorting number. 3) Output structure: Output the initial priority list in CSV format, with the fields: field_pair_id, conflict_rate, distribution_delta, time_overlap, total_score, rank, and store it in the local path / data / initial_priority_list.csv for subsequent modules to read.

[0047] Evaluate the credibility deviation of the source data in each field pair based on the initial field priority list and the field source weight matrix, and perform weighted sorting on the initial field priority list according to the credibility deviation to form a fusion priority vector; In this embodiment, the field pairs are re-sorted based on the field source weight matrix and the initial priority list. The process is as follows: 1) Read the source weight matrix: Each field in this matrix corresponds to a credibility value, and the value range is [0, 1], which is statistically obtained based on the data integrity, update time interval, and access frequency of the field source; 2) Calculate the credibility deviation: Calculate the difference in credibility ΔW = |W1 - W2| for each field pair; 3) Fusion weighted scoring formula: The final fusion priority score = 0.7 * total_score + 0.3 * ΔW, ensuring that both the conflict degree and the credibility difference are comprehensively considered; 4) Vector generation and sorting: Convert the above results into a fusion priority vector, including fields: field pair number, initial score, credibility difference, final weighted score, sorting number. This fusion priority vector is stored in JSON format, and each element contains the field pair identifier and the comprehensive sorting score for the call of the merging strategy module.

[0048] Perform a merging adaptation analysis on the field pairs in the fusion priority vector, match the merging strategy according to the field value range overlap, update synchronization, and semantic mapping relationship, so as to determine the merging strategy type corresponding to each field pair and generate a merging strategy identifier; In this embodiment, the following three-dimensional feature analysis is performed for each field pair to determine the merging strategy type: 1) Field value range overlap analysis: Extract the sample value sets of the two fields and calculate the ratio of their intersection to union (overlap_rate = |A∩B| / |A∪B|). If the overlap rate > 0.6, it is considered that the field value range has the potential for fusion; 2) Update synchronization determination: Extract the field data update timestamp sequences respectively, and use a moving window (length = 6 hours) to calculate the proportion of records that appear in the same time window for the two fields; A synchronization rate > 70% is regarded as time consistency; 3) Semantic mapping relationship determination: Check the semantic labels of the fields. If the labels are the same and the units match, it is determined as semantic consistency; If the labels come from the same upper category (such as "amount" and "payment amount"), they are semantically similar; 4) Strategy type determination rule: a) All three dimensions are satisfied: classified as direct_merge; b) Update synchronization and semantic consistency are satisfied, and the field value range is low: classified as step_merge; c) Only semantically similar, and the others are not satisfied: classified as deferred_merge; d) None of the three are satisfied: classified as reject_merge. Finally, a merge_strategy field is assigned to each pair of fields, and a merging strategy identifier table (fields: field_pair_id, merge_strategy, overlap_score, sync_score, semantic_label_match) is generated and output in CSV format Arrange the merge strategy identifiers in sequence according to the processing priority corresponding to the merge strategy type, and output the field fusion priority schedule table.

[0049] In this embodiment, the merge strategy type is mapped to a fixed execution priority value, and the following process is executed: 1) Priority assignment rule: direct_merge → priority = 1; step_merge → priority = 2; deferred_merge → priority = 3; reject_merge → does not enter the fusion process, and the priority is marked as N / A; 2) Construction of the merge strategy sequence: Sort all merge strategy identifier items in ascending order of the above priorities; 3) Fields in the field fusion priority schedule table: field_pair_id: field pair number; strategy_type: merge strategy type; 4) execution_priority: fusion execution sequence number; 5) instruction_token: instruction corresponding to the preset merge strategy (e.g., merge_union, merge_mapping, merge_skip); This fusion priority schedule table is exported to / plans / final_merge_schedule.csv in tabular form.

[0050] Optionally, the division of the entity semantic level in step S3 is specifically as follows: Set the maximum word length limit to 8, and the word segmentation granularity control parameter to [1, 3]. Perform word segmentation processing and entity recognition on the text description, label field, and interaction record in the service platform integration data. Set the minimum entity frequency threshold to 3 to extract explicit entities, thereby constructing an initial entity set; In this embodiment, word segmentation processing is performed on the text description, label field, and interaction record in the service platform integration data. For this purpose, the maximum word length limit is set to 8, and the word segmentation granularity is controlled by the parameter [1, 3]. The length of each word does not exceed 8 characters, and the granularity is controlled to avoid information loss caused by overly fine segmentation or detail loss caused by overly coarse segmentation. After word segmentation, entity recognition is performed on each text record to extract explicit entities, such as company names, locations, products, etc. The minimum entity frequency threshold is set to 3 to ensure that only entities that appear at least 3 times are extracted as explicit entities. Through this processing, the system can construct an initial entity set and provide a basis for subsequent semantic analysis.

[0051] Perform semantic encoding on each entity node in the initial entity set. Set the maximum input sequence length to 512 and the encoding dimension to 1024. Use a pre-trained context embedding model to extract the entity context vector representation, thereby generating an entity semantic encoding data set; In this embodiment, semantic encoding is performed on each entity node in the initial entity set. Each entity node is encoded through a context embedding model, and the maximum input sequence length is set to 512, that is, the maximum input length of the entity description and its related context is 512 characters. This length is applicable to the context information of the vast majority of entities and ensures the integrity of information. The semantic vector representation of the entity adopts a pre-trained context embedding model with an encoding dimension of 1024, such as BERT or its variants. These models can extract the context information of each entity, thereby generating a high-dimensional entity semantic encoding dataset. Through this process, each entity will be converted into a high-dimensional semantic vector, which can reflect the multiple meanings of the entity in the context.

[0052] Semantic similarity calculation and vector aggregation are performed on the entity semantic encoding dataset. The similarity threshold is set to [0.75, 0.95], and entities with high similarity are classified into the same semantic cluster to generate a semantic aggregation cluster set; In this embodiment, semantic similarity calculation is performed on the entity semantic encoding dataset. By calculating the cosine similarity between entity semantic vectors, the system can judge the similarity between different entities. The similarity threshold is set to [0.75, 0.95], and those entities with similarity within this range are grouped into the same semantic cluster. The setting of the threshold is based on the actual requirements for entity similarity, ensuring that similar entities can be in the same semantic cluster for subsequent semantic aggregation and reasoning. For example, entities with a similarity exceeding 0.95 are considered to have highly similar semantics, while entities with a similarity between 0.75 and 0.95 belong to entities with relatively high similarity. Finally, the system generates a semantic aggregation cluster set, where each cluster contains a group of entities with similar semantics.

[0053] Based on the semantic aggregation cluster set, an entity semantic hierarchical graph structure is constructed; hierarchical division rules are set according to the entity semantic hierarchical graph structure to perform entity hierarchical division, and a semantic hierarchical index table is generated; In this embodiment, each semantic cluster in the semantic aggregation cluster set is used as the initial analysis unit. The semantic tags of the entity nodes therein are extracted, and the clustering center words are extracted. It is set that the length of the entity representation does not exceed 32 characters. When extracting, stop words are excluded, and only the key expressions with the part of speech being a noun or a proper noun are retained. Subsequently, according to the semantic coverage range from large to small, an inclusion relationship index structure between entity aggregation clusters is initially established. For example, the semantic cluster of "Internet companies" may include entity cluster center words such as "A" and "B", so it is determined that "Internet companies" is the upper-level node, forming an initial semantic graph edge structure. In the construction of the graph structure, to ensure the semantic validity of the edges, the minimum semantic association strength threshold is set to 0.8, and only the edges with a significant semantic nesting relationship between the upper and lower entities are retained as the formal graph edges. Each edge records attributes such as direction (upper-level → lower-level), association threshold, and original semantic label matching, and finally a multi-layer semantic hierarchical relationship network graph, that is, an entity semantic hierarchical graph structure, is constructed. After the construction of the graph structure is completed, it enters the hierarchical division stage. In this stage, hierarchical rules are set according to the semantic coverage breadth, vocabulary generality, and upstream and downstream node density: 1) Nodes with a semantic coverage ratio (the ratio of the number of lower-level nodes covered by a certain node to all nodes) ≥ 0.1 are set as the first-level classes; 2) Nodes with an average node depth difference between upstream and downstream greater than or equal to 2 are promoted to the next level; 3) Nodes with a semantic category diversity of the covered child nodes not less than 3 categories (such as "institutions", "persons", "products") can be labeled as generalization classes. According to the above rules, all nodes in the semantic graph are traversed and analyzed to determine their hierarchical labels (such as L1, L2, L3, etc.), and the hierarchical results are output as a semantic hierarchical index table. This table contains fields: entity unique identifier, entity standard name, hierarchical encoding, upper-level node identifier, semantic attribution label, and associated edge weight, etc. The generated index table provides a structured input basis for subsequent entity upper and lower position verification, semantic consistency verification, knowledge path reasoning, etc.

[0054] Perform entity consistency verification and upper and lower position relationship verification on each hierarchical node in the semantic hierarchical index table, remove nodes with semantic ambiguity, and reconstruct the hierarchical path graph to generate entity hierarchical data.

[0055] In this embodiment, entity consistency verification and upper and lower position relationship verification are performed on each hierarchical node in the semantic hierarchical index table. Entity consistency verification aims to ensure that there are no conflicts or ambiguities in the semantics of entities at the same level. For example, it is judged whether two entities representing similar concepts should be merged into the same entity. Upper and lower position relationship verification is used to ensure that the entity relationships between different levels are correct. For example, it is verified whether the upper-level entity actually includes the lower-level entity. For those nodes with semantic ambiguity, they are removed, and the hierarchical path graph is reconstructed if necessary to eliminate inconsistencies. Finally, entity hierarchical data is generated to ensure that each entity correctly reflects its upper and lower position relationships in the hierarchical structure and can provide accurate semantic information for subsequent applications.

[0056] Optionally, the entity relationship reasoning in step S3 is specifically as follows: Perform upper and lower dependency annotation on each level of entity pairs in the entity level data, and construct an entity co-occurrence frequency matrix; In this embodiment, for the multi-level entity data after entity semantic level clustering, first, taking the hierarchical structure tree as the index structure, extract the entity set under each node. Divide the service platform interaction data (in the format of structured record logs, with fields including user ID, access time, service content, etc.) by user session, set the window span to 5 records, and use a sliding window method to extract the context entity set, with the window sliding step size of 1 record. Identify the co-occurring entity pairs in each window, and by comparing their paths in the semantic hierarchical structure, filter out the entity pairs with "parent-child relationship" or "adjacent nodes at the same level". During the annotation process, the following rules are used to judge the upper and lower dependencies: If entity A is above entity B in the hierarchical tree and its occurrence frequency is not less than 1.5 times that of entity B, then mark A as the upper position and B as the lower position. This judgment rule is implemented in the system as a Python structured conditional expression and supports multi-threaded concurrent processing to improve processing efficiency. The co-occurrence frequency and dependency type are encoded into a two-dimensional data structure named co_occurrence_matrix, with fields including: entity ID pair (entity_i, entity_j), co-occurrence times, dependency type (upper / lower / equal level), and hierarchical difference L (e.g., L = 1 indicates direct adjacency). This matrix is finally stored in Parquet format to provide structural input for subsequent edge weight reasoning.

[0057] Based on the entity co-occurrence frequency matrix, embed the entity context vector to predict the edge weight and generate a potential edge set of entity relationships; In this embodiment, all entity pairs are extracted from the co-occurrence frequency matrix, and the context vector representations of each entity are extracted with reference to the semantic coding dictionary (previously trained using the RoBERTa-wwm-ext model with a vector dimension of 1024). The two entity vectors are concatenated to form a 2048-dimensional input feature vector, which is fed into a multi-channel scoring module. This module consists of three sub-structures: ① The semantic cosine similarity module, which calculates the semantic overlap degree of the contexts of the two entities; ② The frequency normalization module, which normalizes the co-occurrence times to the interval [0, 1] and then multiplies by the level difference weight (set to 1 / L_diff); ③ The structural coupling module, which calculates the structural consistency score based on the matching degree between the intersection length of the entity paths and the semantic category labels. Finally, the edge weight is calculated by the following weighted function: Wij = 0.4×Sim + 0.35×Norm + 0.25×Struct; where Wij is the strength score of the connection edge between entity i and entity j, Sim is the semantic similarity score, Norm is the normalized co-occurrence frequency, and Struct is the structural coupling strength. The predicted edge weight threshold is set to 0.6, and only the entity pairs with weights greater than this threshold are retained to generate a potential entity relationship edge set. The edge set is organized in JSONL format, and the fields include the starting entity ID, the target entity ID, the predicted weight value, the sub-items of the scores of each module, and the semantic label pair. The model inference in this step is deployed using TensorFlow Serving and executed in parallel in the GPU environment in units of vector batches to ensure the prediction efficiency.

[0058] Perform ternary rule reasoning on the potential edge set of entity relationships to obtain a set of candidate entity relationship triples; In this embodiment, the defined ternary rule library contains 12 basic relationship patterns (such as <superior organization, contains, subordinate unit>, <technical term, belongs to, industry classification>), and each rule is defined in a logical structure, and the fields include the starting semantic label, the ending semantic label, the relationship keyword, the applicable edge weight range (such as 0.7~1.0), etc. This rule library is maintained in YAML format and loaded through in-memory indexing. For each edge in the potential entity edge set, match the semantic labels and predicted weight values of the two end entities in turn to determine whether the rule trigger condition is met. For example, if the label of entity A is "administrative region", the label of entity B is "service item", and W_ij≥0.75, then the rule "<administrative region, provides, service item>" is matched. After successful matching, a candidate triple (subject, predicate, object) is constructed in the form of a structure, and the matching rule ID, the predicted edge weight, and the rule matching score (calculated comprehensively from the label, structure, and weight) are appended. This processing flow is implemented using a custom relationship matching engine (based on rule tree search) and can record debug logs in batches. The generated set of candidate triples is saved in a structured set named triplet_candidates_v1 for the next step of consistency verification and path completion.

[0059] Perform consistency verification and conflict detection on the candidate triple set of entity relationships, eliminate logically conflicting relationships, and probabilistically complete the incomplete paths in the candidate triple set of entity relationships to generate an entity semantic relationship table; In this embodiment, first, a triple conflict detection rule set is constructed, including 15 types of rules such as mutually exclusive label pairs (e.g., <user type, cannot belong to, multiple identity categories>) and path logical conflicts (e.g., an entity has two different unique superior paths). Each rule is implemented as a Boolean logic expression. Batch verification is performed on each candidate triple, and the conflict is marked with conflict=true and written to a log file for manual review. Subsequently, for the missing path structure (i.e., there is a long distance between two entities in the inference graph but no intermediate node), the maximum number of hops is set to 3, and the co-occurrence frequency matrix and the known edge set are called for relay entity search. The search path score is calculated by the average predicted weight of the intermediate edges, and the minimum confidence threshold is set to 0.65. If the constructed new path meets the scoring conditions, it is included in the triple set and marked as "completed path". The finally generated triple set is output in the form of a relationship table, and the fields include: entity A, entity B, relationship type, whether it is completed, matching score, and whether there is a conflict. By setting the conflict detection backtracking flag bit, duplicate paths can be dynamically filtered out to ensure semantic structure consistency and logical closed-loop.

[0060] Map the entity semantic relationship table to the entity semantic hierarchical graph structure to construct an entity relationship inference graph.

[0061] In this embodiment, based on the previously constructed entity hierarchical graph (the structure is a directed acyclic graph DAG), each triple in the semantic relationship table is mapped to the graph structure in turn. The NetworkX library is used to construct a directed graph object G_semantic_relation, the entity identifier is located at the specific node number in the hierarchical graph, and at the same time the relationship type is mapped to the edge label. The edge weight value (generated by the prediction score) is retained as the graph edge attribute. For entities with multiple path mappings, calculate the structural distance (counted by the number of hops) and semantic weight value of each path, and retain the main path according to the following selection strategy: ; where is the finally selected path with the highest confidence and within the limited number of hops from entity i to entity j in the candidate path set; is the number of intermediate edges or nodes passed from entity i to entity j in path p, that is, the number of hops of the path; is the confidence score for path p; λ1 = 0.6 and λ2 = 0.4 are empirically set. After construction, perform a connectivity check on the entire graph (to determine if there are isolated entities), and use depth-first traversal to detect if there are hierarchical loops. If so, trigger a rollback process. The rollback process means that when structural anomalies (such as hierarchical loops or illegal connections) are found during graph construction, the current round of path mapping and edge insertion operations are automatically revoked to ensure the structural validity and semantic consistency of the final knowledge graph. Specifically, a snapshot of the current graph state is established before each round of path mapping (i.e., backup the current G_semantic_relation object). Once a closed-loop structure is detected (for example, by finding that a visited node is visited again during depth-first traversal, forming a back edge), immediately restore to the snapshot state and revoke all nodes and edges inserted in this round. To improve efficiency, the rollback operation uses the edge difference comparison mechanism of NetworkX, only deleting the newly added edges and the affected path nodes, avoiding repeated construction of stable structures. In addition, for triple relations that frequently cause rollbacks, they are marked as "abnormal candidates" and pushed into the queue to be reviewed for subsequent semantic repair and structural strategy adjustment modules. The output entity relationship inference graph contains node attributes (entity ID, semantic label, hierarchical level) and edge attributes (relationship type, edge weight, path source, completion flag, etc.), which are exported in the Neo4j graph database structure and support subsequent queries and incremental updates.

[0062] Of particular importance is that the triple rule reasoning is specifically as follows: Perform path reachability analysis on each pair of entity nodes in the potential edge set of entity relationships, identify explicit and implicit candidate pairs, and obtain the initial entity pair path set; In this embodiment, for the potential edge set of entity relationships that has been generated in the entity semantic hierarchy graph, the depth-first search (DFS) method is used to perform path reachability analysis on each pair of candidate entity nodes. Set the maximum path hop count Max_Hop = 5, and set the constraint that nodes in the path are not repeated. The path is represented as: p(i,j)={v0 = i, v1, v2,..., v k= j}, where k ≤ Max_Hop; where, i and j are the start and end entities, and v is the intermediate entity node. Each path is attached with the following attribute structure: Hop(p): the number of hops of the path, i.e., the number of edges in the path; Path_Seq(p): the sequence of entity identifiers passed through in the path in turn; IsDirect(p): a boolean flag, set to True if Hop(p) = 1 (explicit path), otherwise False (implicit path); Node_Type(p): the semantic type of each path entity node. After the paths are generated, they are stored as Path_Set = {p(i, j)} ∀(i, j) ∈ E_candidate, and the path sequences, node types, and path identifiers are recorded in a structured table. The obtained initial entity pair path set is used as a structured input for the next stage.

[0063] According to the initial entity pair path set, combined with the entity category rules and relationship pattern rules defined in the knowledge base of the consulting service field, filter the applicable ternary structure rule template set; In this embodiment, based on the path set Path_Set, match the entity category and relationship pattern rules defined in the knowledge base K = {R1, R2,..., R n} of the consulting service field. Each rule template R is formally defined as: R = <T1, rel, T2>, where T1 and T2 are entity semantic categories, and rel is a predefined relationship type (such as "belong to", "govern", "provide"); use the structure matching function Sim_struct(p, R): Sim_struct(p, R) = (∑match(T_node(p), T_R)) / len(p); where T_node(p) is the sequence of types of each node in the path, and T_R is the sequence of node types in the rule template. Set the structure matching threshold θ_match = 0.7. When Sim_struct(p, R) ≥ θ_match, it is considered an adaptable rule template and included in the candidate template set Rule_Candidates(p). The mapping record structure of each pair of paths and their rule templates is as follows: path identifier: p_id; matching template ID: R_id; matching similarity: Sim_struct.

[0064] Match the corresponding rule templates for each entity path in the ternary structure rule template set, and perform relationship composition and logical connection operations to generate candidate relationship inference paths; In this embodiment, for the path set that matches the rule template, logical reduction and relationship synthesis operations are performed according to the template structure. The following synthesis process is used: 1) Node role verification: Ensure that the start point, end point, and intermediate nodes of the path are consistent with the entity categories in the template; 2) Relationship synthesis logic: Combine consecutive semantic edges in the path into a relationship block. If the path is A→B→C and the template is X—Y—Z, the synthesized relationship is rel(A,C); 3) Structure matching rate calculation: Set the matching rate calculation formula: Score(p,R)=|matched_nodes(p,R)| / |total_nodes(R)|; Score(p,R) is the matching rate of the candidate path p and the structure rule template R; matched_nodes(p,R) is the number of nodes in the path p that are exactly the same as the node semantic categories defined in the rule template R. total_nodes(R) represents the total number of nodes defined in the rule template R, that is, the number of entity nodes that should be included in the template structure. Set the synthesis validity threshold θ_struct = 0.85. When Score(p,R)≥θ_struct, an inference path set is generated, and the structure is: <Start_Entity,Relation_Type,End_Entity,Matched_Template_ID,Structure_Confidence>, where Structure_Confidence = Score(p,R) represents the confidence in the consistency between the path and the template structure.

[0065] Extract triples in the form of reasonable structure and closed relationship from the candidate relationship inference paths to generate an initial triple inference set; In this embodiment, in the candidate inference path set, paths that meet the closure and semantic consistency are screened for triple generation. The judgment conditions are as follows: The path structure needs to be single-source and single-endpoint; The maximum number of intermediate entities is 1, and there is a relationship definition for the semantic label in the knowledge base; The relationship edge needs to have a corresponding meaning in the rule template. The paths that meet the above conditions are converted into structured triples: Triple = <Subject,Predicate,Object>; and are attached with the following attributes: Path identifier: p_id; Matching confidence: Confidence(p)=Score(p,R); Relationship semantic source: Rel_Origin(R).

[0066] Perform confidence estimation on the initial triple inference set, and filter out low-confidence relationships to form a candidate triple set of entity relationships.

[0067] In this embodiment, confidence estimation is performed on each triple in the initial triple inference set, and a comprehensive confidence function is set: Conf(Triple)=α·Sim_struct(p,R)+β·Sim_semantic(i,j)+γ·Weight(p); where: Sim_struct(p,R) is the matching degree between the path structure and the template; Sim_semantic(i,j) is the semantic vector similarity between the starting and target entities (calculated by Cosine); Weight(p) is the average edge weight in the path; the weights are set as: α = 0.4, β = 0.4, γ = 0.2. A minimum confidence threshold θ_conf = 0.6 is set, and triples with a value lower than this will be removed from the candidate set. The final retained triple set forms an entity relationship candidate triple set Triple_Set={<s,r,o>|Conf≥θ_conf}, and its path identifier and inference source information are saved.

[0068] Optionally, the window period platform consultation mode recognition in step S4 is specifically as follows: Associate and bind the user consultation preference data with the user entity nodes in the entity relationship inference graph to construct a three-dimensional tensor of user-behavior-time, and generate a candidate time series set for the window period; In this embodiment, first, behavior record fields such as user identifier, consultation timestamp, and involved entity categories (such as "industry planning") are extracted from the platform user consultation preference data, and field matching and unique binding are performed with the user entity nodes already established in the entity relationship inference graph. The lower limit of the entity matching rate is set to 0.85 during this binding process, and only user nodes with highly overlapping matching fields and labels are retained. After the binding is completed, a three-dimensional tensor structure with dimensions of [user ID × behavior type × time step] is constructed, where the time step granularity is set to 12 hours, and the behavior type is divided into 9 types of consultation target labels. According to the time axis of the tensor, the time period is delimited as a 30-day window period, and the behavior sequence data of each user within the window period is extracted, finally forming a candidate time series set for the window period, providing input for subsequent time series clustering processing.

[0069] Perform multi-scale time window segmentation and sliding clustering processing on the candidate time series set for the window period, delimit the consultation active intervals, and perform clustering annotation on the behaviors within each interval to generate an initial behavior window mode set; In this embodiment, for the candidate time series set in the window period, the time window scale set is first set to [6h, 12h, 24h], and the behavior timeline is segmented by multi-scale sliding. The sliding step is set to 2 hours to ensure the ability to capture active behaviors with high time resolution. During the sliding process, the distribution of behavior records within the window is analyzed by density aggregation to identify behavior-intensive intervals, which are determined as user consultation active areas. The behavior data within each active interval is clustered according to the behavior type, the entity category of the consultation object, and the superordinate and subordinate structures, and the clustering granularity is controlled between 5 and 10 categories. Taking the user "U123" as an example, 3 active intervals are formed at the 24-hour scale, mainly clustered by "industry - construction" and "industry - strategy", and finally an initial behavior window pattern set is generated.

[0070] Based on the initial behavior window pattern set, a multi-dimensional behavior spectrogram is constructed, and periodic spectral analysis is performed on the user access category, consultation object, and semantic intention within each window to identify the consultation pattern interval, thereby generating a periodic behavior frequency matrix; In this embodiment, a multi-dimensional behavior spectrogram is first constructed. The axes of the spectrogram include the consultation category frequency, the occurrence frequency of entity semantic labels, and the semantic intention score. The data source is the clustering behavior results in the initial behavior window pattern set. The sliding period length is set to 3 days. By observing the rhythm and intensity changes of the consultation category in different time windows, potential periodic patterns are identified. For each type of behavior in the spectrogram, a Fourier frequency distribution diagram is constructed to identify the frequency peak and mark the period. Taking the example that the frequency of "consultation behavior" appears 3 frequency peaks within a 7-day period, its period label is marked as the 7-day main period. After the period labels are marked for all window behaviors clustering, the frequency data is merged to form a periodic behavior frequency matrix, laying a foundation for subsequent stability analysis.

[0071] Perform a stability score on the periodic behavior frequency matrix for the behavior pattern, screen out the window behavior patterns with high confidence, and generate a list of stable window patterns; In this embodiment, according to the frequency change curve in the periodic behavior frequency matrix, a stability score is performed on each type of periodic behavior. The scoring basis includes the behavior repetition degree, the frequency consistency within the period, and the semantic variance of the behavior intention. The full score of the stability score is set to 1.0, and the behavior patterns with a score lower than 0.65 are not included in the stable pattern library. For example, if a certain type of user consults "Question A" at 8 am every weekday, and this behavior appears in two consecutive periods, the behavior category and entity semantics are consistent, and the semantic intention variance is lower than 0.2, then its stability score is 0.89 and it is retained in the list of stable window patterns. This list contains information such as window ID, behavior category, period label, stability score, etc., and is used as a representative behavior pattern for the subsequent trend fusion process.

[0072] Based on the analysis of the overall consultation behavior trend of the data analysis platform for user consultation preferences, the overall consultation behavior trend data of the platform is obtained; In this embodiment, by aggregating the consultation preference data of all users, the distribution of consultation behavior categories, the trend of consultation time frequency, and the access concentration data of all platform users in the past three months are extracted. The statistical time unit is set to day, and the total frequency of various consultation behaviors is extracted daily and subjected to a moving average process (the moving window is 7 days) to remove accidental interference. In platform trend recognition, key recognition is focused on periodic growth types (such as concentrated consultation on tax filing issues at the beginning of each month), holiday concentration types (such as "consulting on personnel management around Labor Day"), and continuously low-frequency consultation behaviors. The overall consultation trend data of the platform includes three parts: the consultation category time spectrum diagram, the consultation period heat matrix, and the user group behavior density distribution diagram, which are used for aligning and matching stable behavior patterns with the system cycle.

[0073] Align the stable window mode list with the overall consultation behavior trend data of the platform, and identify the system behavior cycle of the stable window mode list to obtain the consultation mode data of the service platform.

[0074] In this embodiment, the stable window mode list and the overall consultation behavior trend data of the platform are subjected to a time axis alignment operation, and the time tolerance range is set to ±12 hours during the alignment process. For each stable behavior pattern, find the matching frequency peak interval in the trend graph. If there is a trend matching area within the tolerance range, it is determined as an effective alignment. Subsequently, the aligned behavior patterns are cyclically modeled according to their occurrence time intervals and the global consultation fluctuation rhythm of the platform to identify their corresponding system behavior cycles. Taking a type of "high-frequency financial and tax consultation" mode as an example, its cycle overlaps with the monthly financial hot spot interval of the platform and is finally marked as a "monthly cycle consultation mode". All behavior patterns that meet cycle overlap and trend consistency are summarized as the consultation mode data of the service platform, and the content includes information such as cycle category, consultation theme, time label, and user group distribution characteristics.

[0075] Optionally, step S5 is specifically: Extract the user historical behavior characteristics from the integrated data of the service platform, and perform user portrait recognition and annotation based on the user historical behavior characteristics to obtain the user portrait label set; In this embodiment, first, the behavior trajectories of users are extracted from the historical behavior logs integrated by the service platform, including fields such as access time, access frequency, access entry path, consultation category, entity label click record, and service feedback record. Users with a time range of the last 90 days and a total number of behavior trajectory records of no less than 100 are included in the scope of portrait analysis. Indicators such as user behavior intensity (daily access frequency), behavior span (number of entity label types covered), and behavior stability (periodic fluctuation range) are used for feature induction, and a behavior intensity threshold of 0.6 and a stability threshold of 0.7 are set for behavior category stability screening. According to the screening results, a preset set of portrait templates is matched, such as "high-frequency technology consultation type", "periodic industry research type", "low-frequency general consultation type", etc., and finally a user portrait label set is generated.

[0076] Based on the service platform consultation mode data and user historical behavior characteristics, a user-scenario-content multi-dimensional behavior tensor is constructed, and the user portrait label set is introduced for multi-dimensional vector expansion to generate a multi-layer perception input vector set; In this embodiment, the extracted user historical behavior characteristics and service platform consultation mode data are structurally integrated. With the user identifier as the primary key, a three-dimensional tensor structure is constructed, and the tensor dimension is set to [user ID × scenario classification × content entity label]. Among them, the scenario classification includes 6 types of context labels such as "topic page access", "historical search", "associated recommendation jump". To enhance the input feature dimension, on the basis of the tensor, the user portrait label set is introduced. Each label is vectorized into a semantic vector with a fixed length (length of 32 dimensions) and spliced to the corresponding position of the user in the tensor to achieve multi-layer perception expansion of behavior data. Taking the user "U456" as an example, the content he accessed covers three entity categories of "consultation", "regional planning", and "technology trend". Combining his portrait labels of "high-frequency industry research type" and "stable access type during a period", an expanded multi-layer perception input vector set with a total dimension length of 256 is generated for subsequent use in the weight building and modeling stage.

[0077] Based on the multi-layer perception input vector set, the behavior paths of users in different scenarios, relationship graphs, and semantic contexts are dynamically weighted and encoded to generate a context behavior perception matrix; In this embodiment, by using the behavior, scenario, and content label dimension data in the multi-layer perception input vector set, and referring to the upper and lower structure and semantic path density of the nodes in the entity relationship graph, a weight model is constructed for each hop behavior in the user access path. In the specific operation, the scenario perception weight is set to 0.4, the graph semantic weight is set to 0.35, and the context label similarity weight is set to 0.25 to calculate the weight of each hop behavior in each behavior path. Taking the path of "industry special page - click on industry label - jump to strategic consulting page" as an example, in the graph, a total of three hop entity nodes of "industry → industry → strategy" are passed through, and the average semantic correlation degree between the nodes is 0.82. Combining the scenario feature annotation and semantic context distance, a final context behavior perception matrix is constructed, and the matrix dimension is [number of behavior hops × number of weight channels].

[0078] Input the context behavior perception matrix into a preset multi-task recommendation representation learning framework, and respectively perform interest preference modeling, real-time intention modeling, and content semantic mapping processing, and fuse the outputs of the three channels to generate a multi-channel recommendation candidate set; In this embodiment, the context behavior perception matrix generated in the previous stage is used as the input and loaded into the multi-task recommendation learning framework. This framework includes three task modules: the interest preference modeling module processes the user's long-term behavior path, identifies the stable preference categories, and sets the analysis period to 30 days; the real-time intention modeling module focuses on the user's current behavior window (the last 48 hours), and sets a weight increase coefficient of 0.2 for the categories with a sudden increase in short-term access frequency; the content semantic mapping module takes the entity label path as the input, constructs the content semantic vector in combination with the edge weight of the relationship graph, and sets the semantic reconstruction distance threshold to 0.75 to filter out abnormal jump behaviors. After the three modules are executed in parallel, the recommendation candidate sets output by each module are merged, and the fusion method adopts a weighted strategy. The default weight distribution ratio is: interest channel 0.5, intention channel 0.3, semantic channel 0.2, and finally a multi-channel recommendation candidate set is generated.

[0079] Perform online collaborative sorting and interpretable fusion evaluation on the multi-channel recommendation candidate set, and output the sorted recommendation content set; In this embodiment, for the generated multi-channel recommendation candidate set, collaborative sorting processing is performed based on the user portrait tag set and the real-time platform hot content. First, according to the semantic tag matching degree score between the user portrait tags and the content tags, a preliminary weight sorting is set. The matching degree scoring standard is that the cosine similarity greater than 0.85 is a high match. Subsequently, interpretability tags are added to each candidate content, such as "recommended due to recently visited content", "recommended due to overlapping with the high-frequency user behavior path", etc. The interpretability score is used as the secondary sorting factor, and the weight ratio is 0.25. Finally, the sorting result is dynamically adjusted according to the user preference trend to generate a sorted recommended content set. Taking the user "U678" as an example, the top 5 contents finally pushed include two hot contents, one strategic content with high historical clicks, one regional dynamic content matching its interest channel behavior, and one real-time news of a highly matched intent channel.

[0080] Perform user strategy matching and platform resource availability comparison on the sorted recommended content set, execute final recommended content filtering, personalized visualization template nesting, and recommendation channel identification coding to generate a dynamic content recommendation data set, and upload it to the consulting service platform for front-end display to execute the content recommendation task.

[0081] In this embodiment, first, each item in the sorted recommended content set is matched item by item with the user's strategy settings. The strategy fields include restrictive conditions such as "avoid duplicate content", "filter records visited in the past 7 days", and "avoid heavy news during non-working hours". After passing the match, further verify the current content availability status in the platform resource scheduling system (including content click-through, service interface online, normal graphic and text loading, etc.), and only retain all available content items. Subsequently, the system calls the personalized visualization template library, selects a recommendation page template according to the visual preference field in the user portrait (such as "graphic card type", "list concise type"), nests the recommended content into the corresponding display module, and respectively adds source coding tags according to the recommendation channels (interest, intent, semantics). The finally generated dynamic content recommendation data set contains fields such as recommendation time, content ID, display template number, and recommendation channel identification, and is uploaded to the platform front-end interface to be displayed in real time on the user recommendation page.

[0082] Optionally, this specification also provides a data integration system for a consulting service platform based on cloud computing, which is used to execute the data integration method for the consulting service platform based on cloud computing as described above. The data integration system for the consulting service platform based on cloud computing includes: A multi-source heterogeneous field fusion module, which is used to obtain multi-source heterogeneous data of the consulting service platform, perform voice and color standardization on the multi-source heterogeneous data of the consulting service platform to obtain voice and color standardized heterogeneous data; perform dynamic adaptation data fusion on the voice and color standardized heterogeneous data to obtain unified adaptation data; The spatio-temporal consistency verification module is used to perform spatio-temporal consistency verification based on the unified adaptation data to obtain the integrated data of the service platform; The entity relationship reasoning module is used to perform entity semantic level division on the integrated data of the service platform to obtain entity level data; perform entity relationship reasoning based on the entity level data to obtain an entity relationship reasoning graph; The consultation mode recognition module is used to analyze the user consultation preference based on the entity relationship reasoning graph, and perform window period platform consultation mode recognition according to the user consultation preference to obtain the service platform consultation mode data; The personalized recommendation module is used to perform multi-layer perceptron personalized recommendation based on the service platform consultation mode data to obtain a dynamic content recommendation data set, and upload it to the consultation service platform to execute the content recommendation task.

[0083] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features invented herein.

Claims

1. A data integration method for a consulting service platform based on cloud computing, characterized in that, It includes the following steps: Step S1: Obtain the multi-source heterogeneous data of the consulting service platform, perform voice sound standardization on the multi-source heterogeneous data of the consulting service platform to obtain voice sound standardized heterogeneous data; perform dynamic adaptation data fusion on the voice sound standardized heterogeneous data to obtain unified adaptation data; Step S2: Perform spatio-temporal consistency verification based on the unified adaptation data to obtain the integrated data of the service platform; Step S3: Perform entity semantic level division on the integrated data of the service platform to obtain entity level data; perform entity relationship reasoning based on the entity level data to obtain an entity relationship reasoning graph; Step S4: Analyze the user's consulting preferences based on the entity relationship reasoning graph, and perform window period platform consulting mode recognition according to the user's consulting preferences to obtain the consulting mode data of the service platform; Step S5: Perform multi-layer perceptron personalized recommendation based on the consulting mode data of the service platform to obtain a dynamic content recommendation data set, and upload it to the consulting service platform to execute the content recommendation task.

2. The data integration method of the cloud computing-based consulting service platform according to claim 1, characterized in that Specifically, Step S1 is as follows: Step S11: Obtain the multi-source heterogeneous data of the consulting service platform, and assign a unique source identifier and structure description metadata to each type of data source to generate an initial heterogeneous data set; Step S12: Perform voice sound standardization on the initial heterogeneous data set to obtain voice sound standardized heterogeneous data, and perform unified field extraction and format normalization processing to generate a structure standardized data set; Step S13: Build a cross-source data mapping matrix based on the structure standardized data set, identify redundant fields and similar semantic units, and perform field merging and priority parsing to generate a semantic mapping data set; Step S14: Combine the source identifier and field priority assignment rules in the semantic mapping data set, perform per-field level data fusion processing, and perform flow control to generate a fusion processing data set; Step S15: Perform outlier removal and keyword field integrity verification on the fusion processing data set to generate unified adaptation data.

3. The data integration method of the cloud computing-based consulting service platform according to claim 2, wherein Specifically, the voice sound standardization in Step S12 is as follows: Extract the audio segments containing voice content in the initial heterogeneous data set, set the sampling rate to 16kHz and mono as the standard input format, and perform resampling and channel unification processing to generate a standardized input audio data set; Perform voice activity detection on the standardized input audio data set, set the silence threshold to -40dB and the minimum voice duration to 300ms, and remove background noise and silent segments to generate a set of effective voice segments; Set the Mel frequency cepstral coefficient dimension to 20, the frame length to 25ms, and the frame shift to 10ms, perform acoustic feature extraction on the set of effective voice segments, and simultaneously extract timbre features including fundamental frequency, fundamental frequency change rate, and formant spacing features to obtain an acoustic feature parameter set; Obtain a regional dialect acoustic feature library, and train a dialect recognition model based on the regional dialect acoustic feature library; Set the frame-level recognition window to 500ms and the sliding step to 250ms, and use the dialect recognition model to perform regional voice feature recognition on the acoustic feature parameter set to generate dialect recognition labels; Obtain a human gender timbre feature library and perform empirical threshold division according to the human gender timbre feature library. The male discrimination threshold range of the fundamental frequency is [85Hz–165Hz], and the female discrimination threshold range of the fundamental frequency is [165Hz–255Hz]; The male discrimination threshold range of the fundamental frequency change rate is [5Hz–15Hz], and the standard deviation of the fundamental frequency change rate < 8Hz; the female discrimination threshold range of the fundamental frequency change rate is [15Hz–35Hz], and the standard deviation of the fundamental frequency change rate > 10Hz; The male discrimination threshold range of the formant spacing is [700Hz–1000Hz]; the female discrimination threshold range of the formant spacing is [1000Hz–1400Hz]; Discriminate the timbre features in the acoustic feature parameter set based on the empirical threshold. When the fundamental frequency, fundamental frequency change rate, and formant spacing simultaneously meet the male / female gender discrimination threshold range, determine the corresponding gender category of the current speech segment and generate a gender recognition label; Perform speech style transfer processing on the acoustic feature parameter set according to the dialect recognition label and gender recognition label, including audio reconstruction of the tone contour in the acoustic feature parameter set according to the dialect recognition label to generate an audio reconstruction speech segment; perform pitch range parameter reconstruction on the timbre features in the acoustic feature parameter set based on the gender recognition label to generate a pitch range reconstruction speech segment; Bind the audio reconstruction speech segment and the pitch range reconstruction speech segment to the standardized input audio data set and semantic label, and replace the original speech field to generate speech timbre standardized heterogeneous data.

4. The data integration method of the cloud computing-based consulting service platform according to claim 3, characterized in that Step S134 is specifically as follows: Step S131: Perform a field-level index scan on the structure-standardized data set, extract the field name, type, value range, and occurrence frequency information of each data source, so as to construct a field attribute feature matrix; Step S132: Aggregate the semantic and structural features between fields based on the field attribute feature matrix, construct a field semantic similarity matrix, and set a semantic fusion threshold to filter out a set of candidate field pairs; Step S133: According to the set of candidate field pairs, establish a cross-source field matching mapping table, and combine the source identifier to construct a two-way mapping graph between fields, so as to generate a cross-source data mapping matrix; Step S134: Perform field conflict analysis on the cross-source data mapping matrix, and perform field priority parsing and merging according to the field conflict analysis result to obtain a field fusion candidate table; Step S135: Perform consistency verification on the field fusion candidate table, eliminate field pairs with a merger conflict rate higher than the set threshold, and perform field-level merging on the remaining field pairs in the field fusion candidate table to generate a semantic mapping data set.

5. The data integration method of the cloud computing-based consulting service platform according to claim 4, characterized in that, Step S134 is specifically as follows: Set the sampling ratio of field value range extraction to 0.8, perform value range intersection analysis and time dimension consistency detection on each field pair in the cross-source data mapping matrix, and count the conflict frequency, value distribution deviation degree, and time effect overlap coefficient to generate a field conflict evaluation index set; Construct a field conflict feature vector based on the field conflict evaluation index set; Use the field conflict feature vector to classify and mark the field conflict mode, and set the conflict severity scoring range to [0,1] to generate a field conflict level annotation table; Combine the field occurrence frequency and data update cycle of field pairs with different conflict levels in the field conflict level annotation table to calculate the field source trust weight, thereby generating a field source weight matrix; According to the field source weight matrix and field conflict level annotation table, priority sorting and merging strategy analysis are performed on each field pair to form a field fusion priority plan table; The merging instructions in the field fusion priority plan table are applied to the cross-source data mapping matrix, and the merging strategy type and credibility label of each field pair are marked to generate a field fusion candidate table.

6. The data integration method of the cloud computing-based consulting service platform according to claim 1, wherein The entity semantic level division in step S3 is specifically as follows: The maximum word length is set to 8, the word segmentation granularity control parameter is set to [1,3], and the text description, tag field and interaction record in the service platform integrated data are processed by word segmentation and entity recognition. The minimum entity frequency threshold is set to 3 to extract explicit entities, thereby constructing the initial entity set. Semantically encode each entity node in the initial entity set, set the maximum input sequence length to 512, the encoding dimension to 1024, and use the pre-trained context embedding model to extract the entity context vector representation, thereby generating an entity semantic encoding dataset; The semantic similarity calculation and vector aggregation are performed on the entity semantic coding dataset. The similarity threshold is set to [0.75, 0.95] to classify entities with high similarity into the same semantic cluster and generate a semantic aggregation cluster set. Based on semantic aggregation clusters, construct entity semantic hierarchical graph structure; According to the entity semantic hierarchy graph structure, a hierarchy division rule is set to divide the entity hierarchy and generate a semantic hierarchy index table; Perform entity consistency verification and hierarchical relationship verification on each hierarchical node in the semantic hierarchical index table, remove semantically ambiguous nodes, reconstruct the hierarchical path graph, and generate entity hierarchical data.

7. The data integration method of the cloud computing-based consulting service platform according to claim 1, wherein The entity relationship reasoning in step S3 is specifically as follows: Label the entity pairs at each level in the entity hierarchy data with their hierarchical dependencies and construct an entity co-occurrence frequency matrix; Based on the entity co-occurrence frequency matrix, the entity context vector is embedded to predict edge weights and generate potential edge sets of entity relationships; Perform ternary rule reasoning on the potential edge set of entity relations to obtain the candidate triple set of entity relations; Perform consistency verification and conflict detection on the candidate triple sets of entity relationships, eliminate logical conflict relationships, and probabilistically complete incomplete paths in the candidate triple sets of entity relationships to generate an entity semantic relationship table; Map the entity semantic relationship table to the entity semantic hierarchical graph structure and construct the entity relationship reasoning graph.

8. The data integration method of the cloud computing-based consulting service platform according to claim 1, wherein The window period platform consultation mode identification in step S4 is specifically as follows: Associating and binding user consultation preference data with user entity nodes in the entity relationship reasoning graph to construct a three-dimensional tensor of user-behavior-time and generate a set of candidate time series for the window period; Perform multi-scale time window segmentation and sliding clustering processing on the window period candidate time series set, divide the consultation active intervals, cluster and annotate the behaviors in each interval, and generate the initial behavior window pattern set; Based on the initial behavior window mode set, construct a multi-dimensional behavior spectrogram, perform periodic spectrum analysis on the user access categories, consultation objects, and semantic intentions within each window, identify the consultation mode interval, and thus generate a periodic behavior frequency matrix; Perform a behavior pattern stability score on the periodic behavior frequency matrix, filter out the high-confidence window behavior patterns, and generate a list of stable window patterns; Based on the data analysis of the overall consultation behavior of the user consultation preference data platform, obtain the overall consultation behavior trend data of the platform; Align the list of stable window patterns with the overall consultation behavior trend data of the platform, and identify the system behavior cycle of the list of stable window patterns to obtain the consultation mode data of the service platform.

9. The data integration method of the cloud computing-based consulting service platform according to claim 1, characterized in that Step S5 is specifically as follows: Extract the user historical behavior characteristics from the integrated data of the service platform, and perform user portrait recognition and annotation based on the user historical behavior characteristics to obtain a user portrait tag set; Construct a user-scenario-content multi-dimensional behavior tensor based on the consultation mode data of the service platform and the user historical behavior characteristics, and introduce the user portrait tag set for multi-dimensional vector expansion to generate a multi-layer perception input vector set; According to the multi-layer perception input vector set, perform dynamic weight encoding on the behavior paths of the user in different scenarios, relationship graphs, and semantic contexts to generate a context behavior perception matrix; Input the context behavior perception matrix into a preset multi-task recommendation representation learning framework, respectively perform interest preference modeling, real-time intention modeling, and content semantic mapping processing, and fuse the three-channel outputs to generate a multi-channel recommendation candidate set; Perform online collaborative ranking and interpretable fusion evaluation on the multi-channel recommendation candidate set, and output a ranked recommendation content set; Perform user policy matching and platform resource availability comparison on the ranked recommendation content set, perform final recommendation content filtering, personalized visualization template nesting, and recommendation channel identification encoding to generate a dynamic content recommendation data set, and upload it to the consultation service platform for front-end display to perform the content recommendation task.

10. A data integration system for a cloud computing-based consulting service platform, characterized in that, For executing the data integration method of the cloud computing-based consultation service platform as described in claim 1, the data integration system of the cloud computing-based consultation service platform includes: A multi-source heterogeneous field fusion module, configured to obtain multi-source heterogeneous data of the consultation service platform, perform voice and color standardization on the multi-source heterogeneous data of the consultation service platform to obtain voice and color standardized heterogeneous data; perform dynamic adaptation data fusion on the voice and color standardized heterogeneous data to obtain unified adaptation data; A spatio-temporal consistency verification module, configured to perform spatio-temporal consistency verification based on the unified adaptation data to obtain the integrated data of the service platform; An entity relationship reasoning module, configured to perform entity semantic level division on the integrated data of the service platform to obtain entity level data; perform entity relationship reasoning based on the entity level data to obtain an entity relationship reasoning graph; A consultation mode recognition module, configured to analyze the user consultation preference based on the entity relationship reasoning graph, and perform window period platform consultation mode recognition according to the user consultation preference to obtain the consultation mode data of the service platform; A personalized recommendation module, which is used to perform multi-layer perception personalized recommendation based on the consultation mode data of the service platform, obtain a dynamic content recommendation data set, and upload it to the consultation service platform to execute the content recommendation task.

Citation Information

Cited By

  • Large-screen data statistical method based on multi-dimensional configuration

    CN120541126A

  • Digital economic label updating system and method based on big data analysis

    CN120631908A

  • A digital economy label updating system and method based on big data analysis

    CN120631908B

  • Data import and intelligent field matching method for low-code platform

    CN120929651A

  • Method for data import and field intelligent matching for low-code platform

    CN120929651B