Communication data intelligent retrieval method and platform based on semantic analysis
By using a user ID-based associated retrieval permission strategy and semantic completion of user profiles, combined with a domain knowledge graph to construct a hierarchical retrieval feature tree, the problem of singular semantic understanding in communication data retrieval is solved, achieving accurate and secure data positioning and improving retrieval efficiency and accuracy.
Patent Information
- Application Number
- CN202511470268.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Existing communication data retrieval technologies suffer from limited semantic understanding, resulting in low retrieval efficiency and difficulty in accurately locating target data, thus failing to meet the needs of efficient applications.
By binding the associated retrieval permission strategy based on user ID, performing semantic completion based on user profile, constructing a hierarchical retrieval feature tree, combining it with the domain knowledge graph for retrieval task matching, and performing top-down conditional dynamic superposition retrieval with permission strategy as constraint, the final output is a time-series retrieval sequence.
It achieves accurate semantic understanding and orderly and efficient retrieval of communication data, can accurately locate data that matches the user's intent, and improves the accuracy and security of retrieval results.
Smart Images

Figure CN120950541B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, and particularly relates to a communication data intelligent retrieval method and platform based on semantic parsing. BACKGROUND
[0002] In real communication scenarios, efficient retrieval of massive communication data is crucial for business analysis, fault diagnosis and user demand response, and accurate acquisition of target data is a core prerequisite for realizing the above applications. In the prior art, communication data retrieval relies on traditional methods such as keyword matching, which plays a certain role in structured data retrieval scenarios. However, as the scale of communication data expands and the semantic complexity increases, the traditional retrieval technology has limitations when applied, which can easily lead to retrieval deviation due to single semantic understanding, making it difficult to accurately locate target data, and the results obtained are disordered and redundant, which cannot meet the needs of accurate communication data retrieval and efficient application. SUMMARY
[0003] The present application provides a communication data intelligent retrieval method and platform based on semantic parsing, which solves the technical problem of single semantic understanding in traditional communication data retrieval, resulting in low retrieval efficiency and difficulty in accurately locating communication data.
[0004] In a first aspect, the present application provides a communication data intelligent retrieval method based on semantic parsing, which comprises: after receiving an original query text submitted by a user, triggering a binding of a retrieval permission policy based on the user ID; loading a user portrait according to the user ID to perform semantic completion of the original query text, obtaining a plurality of semantically equivalent completion sentences; extracting feature words from the plurality of semantically equivalent completion sentences to obtain a plurality of groups of retrieval feature words; performing semantic association divergence of the plurality of groups of retrieval feature words based on a domain knowledge graph to construct a plurality of hierarchical retrieval feature trees; matching a hierarchical limited retrieval task according to the node weight features of the plurality of hierarchical retrieval feature trees, locating a plurality of distributed retrieval units; driving the plurality of distributed retrieval units to perform a top-down conditional dynamic superposition retrieval of the plurality of hierarchical retrieval feature trees under the constraint of the retrieval permission policy, outputting a plurality of refined result sets; calling the metadata time stamps of the plurality of refined result sets to perform descending arrangement and integration, and outputting a communication data time series retrieval sequence.
[0005] In a second aspect of the present application, a semantic analysis-based intelligent communication data retrieval platform is provided, comprising: an association retrieval permission policy execution module, configured to trigger association retrieval permission policy binding based on a user ID after receiving an original query text submitted by a user; a semantic equivalent completion sentence acquisition module, configured to load a user portrait according to the user ID to perform semantic completion of the original query text, and obtain a plurality of semantic equivalent completion sentences; a retrieval feature word acquisition module, configured to perform feature word extraction on the plurality of semantic equivalent completion sentences to obtain a plurality of sets of retrieval feature words; a hierarchical retrieval feature tree construction module, configured to perform semantic association divergence on the plurality of sets of retrieval feature words based on a domain knowledge graph, to construct a plurality of hierarchical retrieval feature trees; a distributed retrieval unit positioning module, configured to perform hierarchical limited retrieval task matching according to node weight features of the plurality of hierarchical retrieval feature trees, to position a plurality of distributed retrieval units; a refined result set acquisition module, configured to constrain the plurality of distributed retrieval units to perform top-down conditional dynamic superposition retrieval on the plurality of hierarchical retrieval feature trees according to the retrieval permission policy, and output a plurality of refined result sets; and a time-sequenced retrieval sequence acquisition module, configured to call metadata timestamps of the plurality of refined result sets to perform descending order arrangement and integration, and output a time-sequenced retrieval sequence of communication data.
[0006] The one or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0007] In the present application, after receiving an original query text submitted by a user, association retrieval permission policy binding based on a user ID is triggered, a user portrait is loaded according to the user ID to complete the semantic of the original query text to obtain a plurality of equivalent sentences, retrieval feature words are extracted, and hierarchical retrieval feature trees are constructed in combination with a domain knowledge graph, distributed retrieval units are matched, and top-down conditional dynamic superposition retrieval is performed according to a permission policy, metadata timestamps of results are called to perform descending order integration, and thus communication data meeting the user's intention is accurately obtained, making the communication data retrieval result more accurate, safe and easy to use, and achieving the technical effects of accurate semantic understanding of communication data, ordered and efficient retrieval, and accurate positioning of communication data meeting the user's intention. BRIEF DESCRIPTION OF DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0009] Figure 1 is a flowchart of the semantic analysis-based intelligent communication data retrieval method provided by the embodiments of the present application.
[0010] Figure 2 Figure 1 is a structural schematic diagram of a communication data intelligent retrieval platform based on semantic analysis provided by an embodiment of the present application.
[0011] Legend: correlation retrieval permission policy execution module 1, semantic equivalent complete sentence acquisition module 2, retrieval feature word acquisition module 3, hierarchical retrieval feature tree construction module 4, distributed retrieval unit positioning module 5, refined result set acquisition module 6, and time-sequenced retrieval sequence acquisition module 7. DETAILED DESCRIPTION
[0012] The present application provides a communication data intelligent retrieval method and platform based on semantic analysis, which is used to solve the technical problem of low retrieval efficiency and difficulty in accurately positioning communication data caused by single semantic understanding in traditional communication data retrieval.
[0013] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0014] It should be noted that the terms "first", "second", etc. in the specification and the above drawings of the present application are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices.
[0015] Embodiment one, as shown in the communication data intelligent retrieval method based on semantic analysis, wherein the method comprises: Figure 1
[0016] Step A100: After receiving the original query text submitted by the user, trigger the correlation retrieval permission policy binding based on the user ID.
[0017] In the embodiments of the present application, the original query text is the retrieval related text submitted by the user, which is the basic input for subsequent semantic completion, feature word extraction and other operations.
[0018] Specifically, after receiving the original query text of the user, real-time request attributes and the user ID are extracted from the request context, an operation risk label is constructed by associating a security audit library, initial search permissions are matched according to the user identity role and are corrected by degradation in combination with the risk label, real-time risk control adaptation is performed according to the operation environment state, and finally an associated search permission policy is output, so as to realize the binding of the associated search permission policy based on the user ID, and the specific steps are described in detail in A110-A140.
[0019] Step A200: According to the user ID, load the user portrait to perform semantic completion of the original query text, and obtain a plurality of semantically equivalent completion sentences.
[0020] In the embodiment of the application, the user portrait is a real-time user portrait feature obtained by searching the user ID in the user portrait database, and specifically includes historical search features, business preference features and semantic habit features.
[0021] Optionally, the real-time user portrait feature is obtained from the portrait database by taking the user ID as a query key, which is spliced with the original query text into a structured input, and a plurality of semantically equivalent completion sentences corresponding to a plurality of personalized search intent vectors are generated by performing multi-branch semantic completion through a semantic completion model, so as to complete the semantic completion of the original query text based on the user portrait, and the specific steps are described in detail in A210-A220.
[0022] Step A300: Obtain a plurality of sets of search feature words by performing feature word extraction on the plurality of semantically equivalent completion sentences.
[0023] In an embodiment of the application, text preprocessing is first performed on the obtained plurality of semantically equivalent completion sentences to filter stop words and punctuation: a preset stop word table is loaded, which contains words such as and, and, and, which have no actual search meaning, each word in the sentence is traversed, the matched stop words are deleted, and the comma, period and other punctuation marks are removed, to obtain clean sentence text without redundancy interference.
[0024] Then, feature word extraction is performed, and the TF-IDF word frequency weight method is used to extract keywords, the word frequency (TF) of each word in each clean sentence is calculated, and the inverse document frequency (IDF) of the word is calculated, the word weight is obtained by multiplying TF and IDF, and this process is the same as that at step A224-4, and will not be described again. Then, the top 3-5 words are sorted in descending order of weight, and the top 3-5 words are taken as the search feature words of the sentence. The search feature words refer to the words that can accurately represent the core search intent of the sentence, for example, the three days in the search of the base station communication data in the past three days are time class search feature words, the search time range is limited, the base station is an object class search feature word, the search data belongs to the object, and the communication data is a type class search feature word, which limits the search data category, and all are key words reflecting the user search demand.
[0025] Finally, the retrieval feature words are grouped, and the extracted retrieval feature words corresponding to each semantic equivalent completion sentence are grouped into a group. For example, the near 3 days, base station, and communication data of the first semantic equivalent completion sentence are the first group, the near 3 days, core base station, and communication record of the second semantic equivalent completion sentence are the second group, and so on, to complete the acquisition of multiple groups of retrieval feature words.
[0026] Through the steps of text preprocessing to remove interference, TF-IDF to extract core feature words, and grouping according to sentences, the effect of accurately extracting multiple groups of retrieval feature words reflecting the core needs of user retrieval from multiple semantic equivalent completion sentences is achieved.
[0027] Step A400: Based on the domain knowledge graph, the semantic association of the multiple groups of retrieval feature words is diverged to construct multiple hierarchical retrieval feature trees.
[0028] Specifically, the first group of retrieval feature words is linked to the nodes of the domain knowledge graph to construct the original semantic layer retrieval feature, and the semantic association of the first hierarchical tree node set is performed based on the root node of the feature and the first initial hierarchical tree is constructed based on the node set. The first historical recurrence frequency set of the node set is locally called to calculate the first feature word initial weight set, and the weight set is mapped and loaded to the initial tree and the weight threshold pruning is performed according to the preset hierarchical decay rule, and finally the first hierarchical retrieval feature tree is constructed, which realizes the purpose of semantic association of the multiple groups of retrieval feature words based on the domain knowledge graph to construct multiple hierarchical retrieval feature trees. The specific steps are described in detail in A410-A450.
[0029] Step A500: According to the node weight features of the multiple hierarchical retrieval feature trees, the hierarchical limited retrieval task matching is performed to locate multiple distributed retrieval units.
[0030] Specifically, first, the first node weight set of the first hierarchical retrieval feature tree is extracted, the first tree depth and the first weight distribution entropy are calculated based on the weight set, the two are weighted and normalized, the first retrieval difficulty coefficient is output, and the first retrieval task level is obtained according to the retrieval task level, and then the first distributed retrieval unit is matched and located according to the retrieval task level, which realizes the hierarchical limited retrieval task matching according to the node weight features of the multiple hierarchical retrieval feature trees to locate multiple distributed retrieval units. The specific steps are described in detail in A510-A540.
[0031] Step A600: With the retrieval permission policy as a constraint, the multiple distributed retrieval units are driven to perform the top-down conditional dynamic superposition retrieval of the multiple hierarchical retrieval feature trees, and multiple refined result sets are output.
[0032] Specifically, the first distributed retrieval unit initiates an initial query based on the top node feature word of the first hierarchical retrieval feature tree according to the retrieval permission policy, outputs a first layer retrieval result, and then dynamically appends the second layer and third layer node feature words to perform secondary and tertiary queries, and iteratively adds the feature words until the first global retrieval result set is output. Finally, the unauthorized fields are removed according to the retrieval permission policy, and the first refined result set is output, thereby completing the retrieval of the multiple hierarchical retrieval feature trees by the multiple distributed retrieval units. The specific steps are described in detail in A610-A650.
[0033] Step A700: The metadata timestamps of the multiple refined result sets are called and arranged in descending order, and a communication data time-sequenced retrieval sequence is output.
[0034] In the embodiments of the present application, the metadata timestamp corresponds to the generation or storage time of the communication data, which is the core basis for time-sequenced arrangement.
[0035] Specifically, after completing the top-down condition dynamic superposition retrieval of the multiple hierarchical retrieval feature trees by the multiple distributed retrieval units and outputting the multiple refined result sets that meet the retrieval permission policy, the metadata timestamps associated with each refined result set are called, then all the refined result sets carrying the timestamps are sequentially sorted according to the descending order of time from new to old, and finally the sorted result sets are integrated as a whole to ensure that the communication data retrieval results in different refined result sets are logically coherent and connected in a unified time sequence. Ultimately, a structured communication data time-sequenced retrieval sequence with clear time sequence is output, which facilitates users to intuitively obtain the latest communication data retrieval results.
[0036] Further, the method provided in the embodiments of the present application includes the following steps:
[0037] A110: Extracting real-time request attributes and the user ID from the request context of the original query text, wherein the real-time request attributes include user identity role and operation environment state.
[0038] A120: Calling the historical unauthorized operation records of the user ID associated security audit library to construct an operation risk label.
[0039] A130: After matching the initial retrieval permission according to the user identity role, performing dynamic permission downgrade correction of the initial retrieval permission based on the operation risk label, and outputting the corrected retrieval permission.
[0040] A140: Performing real-time risk control adaptation on the corrected retrieval permission according to the operation environment state, and outputting the associated retrieval permission policy.
[0041] In the embodiments of the present application, the security audit library is an encrypted database for storing historical search operation records.
[0042] Specifically, after receiving the original query text submitted by the user, the request context of the original query text is first parsed. The request context includes various types of accompanying information when the user initiates the search operation. For example, in the transmission protocol data packet of the search request, a user ID for identifying the user identity will be carried, and at the same time, identity-related information and environmental information when the user initiates the operation will also be included. According to this, the user identity role can be identified, such as a communication network operation and maintenance personnel, an ordinary individual user, and the like, and the operation environment state, wherein the operation environment state specifically covers the login IP geofencing and the device security authentication level, such as whether the user login IP is within the enterprise preset safe geographic area range, and whether the search device used by the user has passed the double authentication and other security authentication processes. Through the extraction of these information, the basic data support for the subsequent construction of the permission policy is provided.
[0043] Then, after the user ID is extracted, the user ID is associated with the security audit library, and the past search operation records of the user are called from the security audit library, and the historical unauthorized operation records are focused and screened, wherein the historical unauthorized operation records are the operation related records of the user when the user performs the communication data search operation in the past, which exceed the effective search permission at that time. The security audit library will store and update the search behavior data of all users in real time, including the permission range of each search of the user, the actual search content, and whether there is an attempt to access or obtain the communication data beyond the permission of the user. For example, if a user has attempted to search the communication logs of other departments not within his responsibility range for 3 times in the past 6 months, based on these historical unauthorized operation records, the corresponding operation risk label is constructed for the user, and common label types include high risk, medium risk, and low risk.
[0044] Further, the division rules of the above three risk labels are mainly based on the frequency of unauthorized operations, the severity of unauthorized behavior and the potential impact of the user within a preset time range. The low-risk label is suitable for users who have less than 2 unauthorized operations in the past 6 months, and the unauthorized attempts do not involve highly sensitive communication data such as core network topology data, user privacy data, etc., and do not pose a substantial threat to data security. The medium-risk label is for users who have 2-3 unauthorized operations in the past 6 months, or whose unauthorized behavior involves moderately sensitive data such as department-level communication log summaries. Such operations do not lead to data leakage, but there is a clear intention to break the boundaries of authority. The high-risk label is suitable for users who have more than 3 unauthorized operations in the past 6 months, or whose unauthorized behavior involves highly sensitive data such as user call content, encrypted communication protocol data, etc., or whose unauthorized attempts have partially succeeded in obtaining data, such as more than 3 attempts to retrieve other department core communication fault records, or 1 unauthorized success in obtaining sensitive data. If a user meets the conditions of different risk levels at the same time, follow the principle of "high not low", and mark the higher risk level first.
[0045] After completing the operation risk label construction, the corresponding initial search permission is matched according to the extracted user identity role. Specifically, based on the actual business needs and security specifications of communication data management, different user identity roles correspond to different initial search permissions, such as the initial search permission of communication network maintenance personnel is to view all base stations in the area they are responsible for Real-time communication data and historical logs, and the initial search permission of ordinary individual users is only to view their own call records, SMS records and other personal communication data.
[0046] Subsequently, based on the constructed low, medium and high three operation risk labels, the matched initial search permission is executed for differential dynamic degradation correction: for users with low-risk operation risk labels, only slight restrictions are made on the basis of initial search permissions, usually a small amount of non-core communication data search range is narrowed, or the historical data backtracking time is adjusted, etc. The edge permissions are limited to minimize the impact on the user's normal search needs, for example, an ordinary user whose initial permission is to view his own call records, SMS records and other personal communication data in the past 3 months, the system will slightly degrade his permission and adjust it to only allow him to view his own call records and SMS records in the past 1 month, reducing unnecessary data access range.
[0047] Then for the user with medium risk operation risk label, the initial search permission is moderately contracted, which is specifically represented as narrowing the search range of core communication data or limiting key search functions such as data export and detail viewing, while ensuring basic use and strengthening risk control. For example, the initial permission of the operation and maintenance personnel who is responsible for the communication data of 5 base stations in the responsible area is to view the basic communication logs of the 5 base stations. The system moderately reduces the permission of the operation and maintenance personnel and adjusts the permission to allow the operation and maintenance personnel to view the basic communication logs of only 3 main responsible base stations, thereby executing the modification of reducing the data access scale. For the user with high risk operation risk label, the initial search permission is strictly limited, which is specifically represented as greatly narrowing the search range of core communication data or closing important search permissions such as real-time data viewing and sensitive field access, thereby significantly reducing the permission level to maximize the avoidance of the risk of unauthorized operation. For example, the initial permission of the operation and maintenance personnel who is responsible for the communication data of all base stations in the responsible area is to view the communication data of the 3 specific base stations directly responsible for by the operation and maintenance personnel. The system greatly reduces the permission of the operation and maintenance personnel and adjusts the permission to allow the operation and maintenance personnel to view the communication data of only the 3 specific base stations, thereby strictly limiting the data access boundary. Finally, the modified search permission is executed in real time according to the operation environment state extracted by the above steps. For example, the modified search permission of a user is to view the communication data of 3 specific base stations. At this time, the system checks the operation environment state of the user. If it is found that the user logs in from a public network outside the preset safe geographic fence, that is, the user usually logs in to the system in the company intranet, or the security authentication level of the device used by the user is low, such as the device does not enable dual authentication and it is detected that the device has an unpatched security vulnerability, the modified search permission is further adjusted, for example, the user is limited to view only the communication data of the 3 base stations within 24 hours, instead of all historical data. After the above series of adaptive adjustments, the final associated search permission policy is output, which clearly indicates the communication data range, data type and operation restrictions that the user can access during the current search operation.
[0048] Finally, the modified search permission is executed in real time according to the operation environment state extracted by the above steps. For example, the modified search permission of a user is to view the communication data of 3 specific base stations. At this time, the system checks the operation environment state of the user. If it is found that the user logs in from a public network outside the preset safe geographic fence, that is, the user usually logs in to the system in the company intranet, or the security authentication level of the device used by the user is low, such as the device does not enable dual authentication and it is detected that the device has an unpatched security vulnerability, the modified search permission is further adjusted, for example, the user is limited to view only the communication data of the 3 base stations within 24 hours, instead of all historical data. After the above series of adaptive adjustments, the final associated search permission policy is output, which clearly indicates the communication data range, data type and operation restrictions that the user can access during the current search operation.
[0049] Through the consecutive steps of extracting multi-dimensional information from the request context, constructing a risk label by associating a security audit library, dynamically adjusting the permission in combination with the identity and the risk, and executing real-time risk control adaptation according to the operation environment, the effect of accurately binding the associated search permission policy for the user is achieved, while ensuring the normal search demand of the legal user, effectively preventing unauthorized operation and protecting the security of communication data.
[0050] Further, step A200 in the method provided by the embodiment of the application comprises:
[0051] A210: query the user portrait database with the user ID as a query key to retrieve real-time user portrait features.
[0052] A220: The original query text and real-time user portrait features are spliced as structured input, multi-branch semantic completion is performed via a semantic completion model, and the plurality of semantic equivalent completion sentences corresponding to the plurality of personalized retrieval intention vectors are generated.
[0053] Optionally, in the process of loading the user portrait according to the user ID, after receiving the original query text submitted by the user, the user ID is taken as a core query key to access a pre-constructed user portrait database, which is a structured data set for storing user personalized features, and is internally uniquely indexed by the user ID to store dynamic data related to user retrieval behavior, business attributes and semantic habits.
[0054] Then, based on the uniqueness of the user ID, the portrait data entry corresponding to the user is quickly located in the index system of the database, and real-time user portrait features are retrieved and extracted therefrom, wherein the real-time user portrait features are not static data, but a multi-dimensional feature set dynamically updated with user behavior, specifically including historical retrieval features, business preference features and semantic habit features, which are described in detail in step A211.
[0055] Next, the real-time user portrait features are converted into a multi-dimensional dynamic feature vector with dynamic free attributes, the query intention sequence is generated by extracting the word segmentation from the original query text and inserting the intention anchor point symbol, and then the multi-dimensional dynamic feature vector is updated by random disturbance based on the probability distribution of the sequence to splice and construct a plurality of personalized retrieval intention vectors. Finally, the original query text and the personalized retrieval intention vectors are taken as structured input, multi-branch semantic completion is performed via a semantic completion model, and a plurality of semantic equivalent completion sentences are generated, which are described in detail in steps A221-A224.
[0056] Further, the method provided in the embodiment of the application comprises the following steps:
[0057] A211: The real-time user portrait features include historical retrieval features, business preference features and semantic habit features.
[0058] Specifically, the historical retrieval features are user past retrieval behavior data recorded in the user portrait database, specifically covering high-frequency retrieval keywords, i.e. core words repeatedly used by the user in multiple retrievals; commonly used filtering condition combinations, i.e. conditions habitually combined by the user when performing retrieval, such as setting the time range to the last 7 days and the data type to real-time communication traffic; repeated query patterns, i.e. periodic or repetitive retrieval behavior rules of the user, such as fixedly retrieving department voice call statistical data of the last week every Monday.
[0059] The business preference feature reflects the retrieval tendency of a user based on the user's business needs or job responsibilities, including frequently accessed communication data types, data categories that the user prefers to search, such as focusing on voice call records, SMS content, or enterprise email communication data; and key contacts / departments that the user focuses on, that is, specific objects that the user frequently associates when searching, for example, a customer service personnel often focuses on the communication records of the complaint handling department or core customers.
[0060] The semantic habit feature reflects the individualized language habits of a user in retrieval expressions, and the core is an individualized term table, that is, the exclusive expressions that the user is used to use when searching, for example, most users in daily life use the term "voice" to refer to the technical level of VoIP voice communication, and use the term "message" to refer to short message service data. Through such retrieval and extraction processes, real-time user portrait features that can comprehensively and accurately depict the retrieval demand tendencies of a user are obtained, and specific and user demand-oriented individualized data bases are provided for subsequent steps, so as to ensure that the subsequent semantic completion operation can accurately match the potential retrieval intention of a user.
[0061] Further, the method provided in the embodiment of the present application comprises the following steps A220:
[0062] A221: converting the real-time user portrait feature into a multi-dimensional dynamic feature vector, wherein the multi-dimensional dynamic feature vector has a dynamic free attribute.
[0063] A222: extracting from the original query text, inserting an intent anchor point symbol after segmentation, and generating a query intent sequence.
[0064] A223: according to the query intent sequence, after updating the vector weight random disturbance based on the probability distribution of the multi-dimensional dynamic feature vector, the multi-dimensional dynamic feature vector is constructed by splicing vectors.
[0065] A224: taking the original query text and the plurality of individualized retrieval intent vectors as structured inputs, performing multi-branch semantic completion via the semantic completion model, and generating the plurality of semantically equivalent completion sentences.
[0066] Specifically, first, the real-time user portrait features are processed by vector conversion. The real-time user portrait features obtained in the foregoing steps are called, which include three types of core dimension data, i.e., historical search features, business preference features, and semantic habit features. The feature data is processed by standardization, and then converted into a multi-dimensional dynamic feature vector by word embedding and feature embedding technology. The vector is one-to-one corresponding to the categories of real-time user portrait features in terms of the dimension of the vector. The numerical value of each dimension represents the importance of the corresponding feature. The dynamic free attributes contained in the vector mean that the weight of each dimension in the vector is not fixed and will be dynamically adjusted according to the changes of the user's recent search behavior and the differences of the current search scene, so as to ensure that the vector can reflect the user's search demand tendency in real time. The specific process is as follows:
[0067] Step a: When the real-time user portrait features are processed by standardization, corresponding standardization methods are used for different types of feature data to ensure that the feature data format is uniform and the numerical range is adapted to the subsequent embedding process. For the numerical features in the historical search features, such as search frequency and use frequency, the Min-Max normalization method is used: first, the maximum and minimum values of the numerical features in the user portrait database are obtained, and the original numerical value is mapped to the [0, 1] interval by the formula normalized value=(original value-min value) / (max value-min value), so as to eliminate the influence of different numerical magnitudes on the subsequent vector weight; for the classification features in the business preference features, such as the commonly accessed communication data types including voice, SMS, and email, the key departments to be focused on including the operation and maintenance department and the customer service department, and the classification data in the semantic habit features, such as the technical classification corresponding to the personalized terms, a separate binary feature bit is generated for each classification category by using one-hot encoding. When the user feature belongs to a certain category, the corresponding feature bit is assigned a value of 1, and the rest is assigned a value of 0. For example, if the user commonly accesses voice communication data, the feature bit corresponding to voice is 1, and the feature bits corresponding to SMS and email are 0, so as to avoid the disorder of classification data interfering with feature embedding.
[0068] Step b: When converting the normalized feature data into a multi-dimensional dynamic feature vector, the method of fixed dimension corresponding feature category + direct weight assignment can be used: First, set the total vector dimension according to the three core dimensions of the real-time user portrait features, including historical search features, business preference features, and semantic habit features. For example, set 3 dimensions for historical search features, corresponding to high-frequency search keywords, commonly used filtering condition combinations, and repeated query patterns; set 2 dimensions for business preference features, corresponding to commonly accessed communication data types and focus objects; and set 2 dimensions for semantic habit features, corresponding to personalized terms and commonly used expression styles, forming a 7-dimensional vector. Then, assign weights to each dimension: normalized numerical features can be directly used as corresponding dimension weights, and normalized categorical features are assigned a weight of 0.9 for matching category dimensions and a weight of 0.1 for non-matching category dimensions. Through the method of dimension corresponding features + direct filling, a multi-dimensional dynamic feature vector is quickly formed, and the dynamic nature can be reflected by adjusting the weights of each dimension as the user's features change.
[0069] Step c: When realizing the dynamic free attribute of the multi-dimensional dynamic feature vector, the recent behavior statistics + scene factor adaptation method is used based on the obtained real-time user portrait features: First, regularly count the user's search behavior data in the past 7 days, such as the number of times the high-frequency keywords in the historical search features are used and the access frequency of each type of communication data in the business preference features. For high-frequency feature dimensions, increase the weight by 0.1 for every 1 more high-frequency use, with a maximum weight of 0.9, and decrease the weight by 0.05 for every 1 less use, with a minimum weight of 0.1. Then extract the key information of the current search scene, such as whether the login device is an office computer or a personal mobile phone, and whether the search time is during work hours or non-work hours. The preset scene factors are 1.0 for office computers and work hours, and 0.7 for personal mobile phones and non-work hours. Multiply the adjusted weights by the corresponding scene factors to get the final dynamic weights, so that the weights of each dimension of the vector change with the user's recent search behavior and the current scene, thus realizing the dynamic free attribute.
[0070] Next, after completing the construction of the multi-dimensional dynamic feature vector, the original query text is structured. Specifically: first, the original query text submitted by the user is segmented by words, and the continuous text is split into independent semantic units, i.e. the word segmentation result, to ensure that the core information in the text can be accurately extracted. Subsequently, an intent anchor symbol is inserted in the word segmentation result. The intent anchor symbol is a symbol used to mark the core retrieval intent in the word segmentation result. It can be divided into time anchor symbol, object anchor symbol, action anchor symbol, etc. corresponding to the time information, retrieval object information, operation action information in the text. Its role is to provide clear intent guidance for subsequent vector weight adjustment. For example, if the original query text is to retrieve the core department voice communication record in the past 3 days, the word segmentation result is to retrieve, the past 3 days, the core department, and the voice communication record. The action anchor symbol can be inserted at retrieve, the time anchor symbol can be inserted at the past 3 days, the object anchor symbol can be inserted at the core department, and the object anchor symbol can be inserted at the voice communication record. Finally, the query intent sequence containing the anchor symbol mark is generated.
[0071] After that, the multi-dimensional dynamic feature vector is updated and the vector splicing operation is performed according to the query intent sequence. First, the generated query intent sequence is loaded, and the core information marked by the intent anchor symbol in it is parsed, such as the past 3 days corresponding to the time anchor symbol in the above example, the voice communication record corresponding to the object anchor symbol, etc. The multi-dimensional dynamic feature vector dimension associated with these core information is determined. Then, a vector weight random disturbance update method based on probability distribution is used. According to the core information of the query intent sequence, the vector weight of the associated dimension is randomly adjusted within the preset probability distribution range, such as normal distribution. This adjustment is guided by the information marked by the intent anchor symbol, ensuring that the adjusted weight can better fit the current retrieval intent. For example, if the query intent sequence contains the object anchor symbol at the voice communication record, the system will randomly disturb and update the weight of the communication data type dimension corresponding to the business preference feature in the multi-dimensional dynamic feature vector within the normal distribution range. Finally, the vectors combined by different weights obtained after multiple disturbance updates are spliced to generate multiple differentiated personalized retrieval intent vectors, each vector corresponding to a user's potential retrieval intent direction.
[0072] Further, in determining the normal distribution range, the initial weight of the associated dimension in the multi-dimensional dynamic feature vector (such as the initial weight of 0.9 in the matching category and 0.1 in the non-matching category in step b above) and the closeness of the associated dimension to the search intent can be combined to set: if the core information of the intent anchor symbol is strongly related to a certain dimension, the mean of the normal distribution can be set to 0 and the standard deviation to 0.1, to ensure that the adjustment range is moderate and does not deviate from the core range of the initial weight; if the association is weak, the standard deviation can be reduced to 0.05 to avoid excessive disturbance. In specific operation, first, the associated dimension is locked according to the intent anchor symbol, for example, the time anchor symbol corresponds to the repeated query mode dimension of the historical search feature, a random number is generated within the set normal distribution, the random number is added to the current weight of the dimension, and finally the result is interval truncated, such as being limited to 0-1, to prevent the weight from exceeding the effective range, that is, the intent-oriented random disturbance update is completed, and the updated weight is ensured to be more consistent with the current search intent.
[0073] Finally, the method constructs a first intent identifier based on the first personalized search intent vector and splices the first intent identifier with the original query text to generate a first structured input matrix, drives the semantic completion model to perform a beam width adjustable search decoding on the matrix to output a first semantic equivalent candidate set, extracts a time entity and a core action from the original query text to construct a core entity conservation rule, and finally calculates the semantic similarity between the candidate set and the original query text under the constraint of the rule to filter out a first semantic equivalent completion sentence. Steps A224-1 to A224-4 are described in detail.
[0074] Through the above steps, the user's personalized features are deeply integrated with the current search intent, providing multi-directional and accurate input for the subsequent semantic completion model, and thus covering multiple potential search intents of the user.
[0075] Further, the method provided in the embodiments of the present application includes the following steps A224.
[0076] A224-1: Constructing a first intent identifier based on the first personalized search intent vector, and splicing the first intent identifier with the original query text to generate a first structured input matrix.
[0077] A224-2: Driving the semantic completion model to perform a beam width adjustable search decoding on the first structured input matrix to output a first semantic equivalent candidate set.
[0078] A224-3: Extracting a time entity and a core action from the original query text to construct a core entity conservation rule.
[0079] A224-4: constrained by the core entity conservation rule, the semantic similarity calculation of the first semantic equivalent candidate set and the original query text is performed to obtain the first semantic equivalent completion sentence.
[0080] Specifically, first, a first intention identifier is constructed based on a first personalized search intention vector. Any personalized search intention vector reflecting the user's specific search tendency is converted into a fixed-dimension dense vector as the first intention identifier through Word2Vec technology. The identifier explicitly identifies the structured identification of the semantic completion direction, avoiding model deviation from the user's personalized needs. Then, the first intention identifier is spliced with the segmentation vector of the original query text by row or column, where the segmentation vector is obtained by processing the original query text through the jieba segmentation tool, and has the same numerical vector format as the first intention identifier. After splicing, a first structured input matrix is generated to provide structured input data for the semantic completion model.
[0081] Then, the semantic completion model is driven to perform a beam width adjustable search decoding on the matrix. The semantic completion model is constructed using a unidirectional LSTM neural network model: the model input layer receives the first structured input matrix, and converts the matrix data into a sequence format that can be processed by the model; the hidden layer captures the text semantic association in the input sequence layer by layer through the unidirectional LSTM network, such as the association between the original query text and the user's personalized intention; the output layer outputs the probability distribution of each position that can generate a word through a fully connected layer and a softmax activation function. During the training process, a paired data set of search queries-semantic completion sentences in the communication field is used, such as search base station fault data→search 5G base station fault communication data in the past 7 days, etc. The model parameters are iteratively optimized using a cross-entropy loss function, so that the model gradually acquires the ability to generate semantically equivalent sentences. In the search decoding stage, a beam search method is used, and the beam width is set to 3, i.e. the top 3 candidate fragments with the highest semantic scores are retained each time, and the generation process is dynamically adjusted, and finally a first semantic equivalent candidate set containing multiple candidate sentences is output.
[0082] Subsequently, the core entity conservation rule is constructed by extracting the time entity and the core action from the original query text. The BERT-based named entity recognition method is used to identify the time entity from the original query text, which is the time range defined in the search request and is a key feature to ensure the timeliness of the search results, such as the past 3 days, a certain year and month, etc. The core action refers to the operation behavior expected by the user, which is a core element to define the type of search task, such as search, extraction, etc. Then, the core entity conservation rule is constructed, i.e. the completion sentence selected subsequently must completely retain these time entities and core actions, and cannot be replaced or deleted at will, to ensure that the completion does not deviate from the user's original search intention.
[0083] After that, each candidate sentence in the first set of semantically equivalent candidates and the original query text are respectively converted into a multi-dimensional semantic vector using the TF-IDF word frequency weight method + basic segmentation method: first, use a conventional segmentation tool (such as jieba) to segment each candidate sentence in the first set of semantically equivalent candidates and the original query text respectively, and remove stop words that have no actual semantics, such as and the like; then, count all segmented words to construct a unified word table containing all unique words in the text; then, calculate the TF-IDF value of each word in the text, where TF represents the proportion of the number of occurrences of the word in the current text, TF = the number of occurrences of a word in the current text ÷ the total number of words in the current text, and IDF represents the scarcity of the word in all texts, IDF = log 10 (total number of texts ÷ number of texts containing the word), where the fewer the number of occurrences of a word in all texts, the more scarce it is, the greater the IDF value, the higher the distinguishing degree of the word to the text, which can more accurately distinguish the semantic differences between different texts, and then the product of the two is the weight of the word; finally, arrange the TF-IDF values of all words in the word table corresponding to each text in order to form a numerical sequence, which is the multi-dimensional semantic vector of the text.
[0084] Finally, the core entity conservation rule is used as a constraint to select the completed sentence using the cosine similarity calculation method. The semantic similarity between the two is calculated by cosine similarity = (vector A • vector B) / (||vector A|| * ||vector B||), with a value range of [0, 1], and the closer the value to 1, the more similar the semantics. The similarity threshold is set to 0.8, and the candidate sentence that meets the core entity conservation rule and has a semantic similarity of ≥0.8 is selected, i.e., the first semantically equivalent completed sentence is obtained.
[0085] Through the above steps, combined with the semantic completion model trained by communication field data, the effect of generating accurate and semantically equivalent completed sentences while strictly preserving the core elements of the user's original search intent is achieved, providing high-quality input for subsequent multi-group search feature word extraction.
[0086] Further, the method provided in the embodiments of the present application comprises the following steps A400:
[0087] A410: constructing an original semantic layer search feature by performing entity linking on the first group of search feature words and the nodes of the domain knowledge graph.
[0088] A420: taking the original semantic layer search feature as a root node, performing semantic association divergence of a preset extension level to obtain a first level tree node set.
[0089] A430: constructing a first initial hierarchical tree based on the first level tree node set.
[0090] A440: locally call the first historical recurrence frequency set of the first hierarchical tree node set, and calculate a first feature word initial weight set based on the first historical recurrence frequency set.
[0091] A450: after mapping and loading the first feature word initial weight set to the first initial hierarchical tree, perform weight threshold pruning based on a preset hierarchical attenuation rule to construct a first hierarchical retrieval feature tree.
[0092] In an embodiment of the present application, the domain knowledge graph refers to a special knowledge graph suitable for a communication data retrieval scenario, which contains nodes such as communication equipment, data types, time dimensions, and the associated relationships between the nodes, and can be directly reused by those skilled in the art.
[0093] In one embodiment, first, a first set of retrieval feature words extracted from the plurality of semantically equivalent completion sentences in the foregoing steps, i.e., a set of words that can accurately represent the core needs of user retrieval, are matched and associated with the nodes of the domain knowledge graph. First, string similarity matching is performed to calculate the similarity between the base station and the communication infrastructure-base station node name in the domain knowledge graph, for example, the edit distance + substring matching supplement method can be used, and the process is as follows:
[0094] First, the characters of the two strings are unified in size and case, and no additional processing is required for excess symbols; then, the minimum number of editing operations required to convert the base station to the communication infrastructure-base station is calculated, which is only 7 insertion operations by inserting the communication infrastructure- before the base station, and the editing number is 7; then, based on the length of the longer string communication infrastructure-base station, which is 9 characters, the base similarity is calculated as 1-editing number / length of longer string = 0.22, and the entity linking is completed by similarity calculation. Then, combined with the node attributes such as the alias and the classification of the node, it is determined that there is a unique correspondence between the first set of retrieval feature words and the nodes of the domain knowledge graph, and finally the associated node group is combined as the original semantic layer retrieval feature, so that the originally isolated retrieval feature words have a domain semantic association basis.
[0095] Next, taking the original semantic layer retrieval feature as the root node, the preset extension level is traversed and executed using the knowledge graph association relationship, such as the preset semantic association divergence of 2 layers of extension: based on the hierarchical relationship and attribute association relationship between the nodes in the domain knowledge graph, the nodes are expanded layer by layer, for example, taking the root node of the base station as the starting point, the first layer diverges the hierarchical relationship of its lower nodes 5G base station and 4G base station, and the second layer further diverges the attribute associated nodes 5G macro base station and 5G micro base station from the 5G base station. Finally, all the diverged nodes form a first hierarchical tree node set, which contains the root node and the retrieval feature words of the semantic association of each level, realizing the semantic range expansion of the retrieval feature words.
[0096] Subsequently, based on the first hierarchical tree node set, a first initial hierarchical tree is constructed using a recursive construction method: the original semantic layer retrieval feature is set as the top root node of the tree, the 5G base station and the 4G base station obtained by the first layer divergence are set as the child nodes of the root node, i.e., the second layer, the 5G macro base station and the 5G micro base station obtained by the second layer divergence are set as the child nodes of the 5G base station, i.e., the third layer, forming the first initial hierarchical tree with a clear hierarchical structure. At this time, only node information is contained in the tree, and the retrieval importance of the node is not embodied.
[0097] Then, the first historical recurrence frequency set of the first hierarchical tree node set is locally called, i.e., the number of times each node feature word is used as a retrieval condition in the past retrieval process stored locally by the system is recorded. For example, the 5G base station is used 150 times in the past three months, and the 4G base station is used 90 times. The first feature word initial weight set is calculated using a frequency proportion weighting method: first, the total historical use frequency of all nodes in the first hierarchical tree node set is counted, which is assumed to be 300 times, and then the use frequency proportion of each node is calculated. For example, the weight of the 5G base station is 150 / 300 = 0.5, and the weight of the 4G base station is 90 / 300 = 0.3. The weight value directly reflects the retrieval heat of the node feature word. The higher the heat, the greater the weight, and the node is preferentially matched in subsequent retrieval.
[0098] Finally, the first feature word initial weight set is mapped and loaded to the first initial hierarchical tree according to the node correspondence relationship, and weight threshold pruning is performed using hierarchical decay threshold pruning: the preset hierarchical decay rule is that the node weight needs to be multiplied by a decay coefficient of 0.8 for each next level, and the minimum retention threshold of the weight threshold is 0.2. For example, the initial weight of the 5G macro base station in the third layer is 0.5 x 0.8 = 0.4. Since it extends one level from the 5G base station in the second layer, the decay operation needs to be performed. If the weight is higher than the minimum retention threshold 0.2, it is retained. Nodes with a weight lower than 0.2 are pruned and removed. Finally, the first hierarchical retrieval feature tree with a simplified structure, reasonable node weight, and clear hierarchical semantics is constructed.
[0099] Through the above steps and in combination with the knowledge graph in the communication field and the historical retrieval data, the isolated retrieval feature word is converted into a hierarchical retrieval feature tree with close semantic association, weight adaptation to retrieval requirements, and a simplified structure, thereby providing efficient feature support for subsequent accurate retrieval.
[0100] Further, the step A500 in the method provided by the embodiment of the application includes:
[0101] A510: Extracting a first node weight set of the first hierarchical retrieval feature tree.
[0102] A520: Calculating a first tree depth and a first weight distribution entropy based on the first node weight set.
[0103] A530: After the first search difficulty coefficient is output by weighting and normalizing the first tree depth and the first weight distribution entropy, a first search task level is output according to the first search difficulty coefficient.
[0104] A540: A first distributed search unit is matched according to the first search task level.
[0105] In the embodiments of the present application, the first distributed search unit is a search unit adopting a distributed search node architecture, such as a Spark-based distributed computing node, and has independent data query and processing capability.
[0106] Optionally, first, for the constructed first-level search feature tree, a tree node traversal algorithm is used to extract a first node weight set: all nodes of the first-level search feature tree, including the root node and each level of child nodes, are accessed in turn by depth-first traversal or breadth-first traversal, and the initial weight of the feature word preloaded by each node is read, the weight value being derived from the first feature word initial weight set calculated based on the historical recurrence frequency in the foregoing step and pruned by level attenuation; all node weight values are arranged in the form of a set in the order of node levels, that is, the first node weight set is obtained.
[0107] Then, the first tree depth and the first weight distribution entropy are calculated based on the first node weight set. In the calculation of the first tree depth, the longest path node counting method is used to count the total number of nodes of the path containing the most nodes from the root node, and then subtract 1, so that the path length = node number - 1, which is the first tree depth. The depth reflects the complexity of the structure of the search feature tree, and the greater the depth, the more levels need to be traversed during search. In the calculation of the first weight distribution entropy, the information entropy calculation formula H = -∑(pi x log2 pi) is used, where pi is the proportion of the weight of a single node to the total weight of all nodes. The total weight of all weights in the first node weight set is calculated first, and then the proportion pi of each weight is calculated in turn. The first weight distribution entropy is calculated by substituting the formula, and the greater the entropy value, the more uniform the node weight distribution, and the more nodes that need to be considered during search.
[0108] Then, the first tree depth and the first weight distribution entropy obtained are processed by using a weighted normalization method, and a first search difficulty coefficient is output, and the coefficient is used for task classification. Specifically, first, the first tree depth and the first weight distribution entropy are normalized respectively, and the values are mapped to the interval [0, 1]. For example, if the preset maximum tree depth is 5, and the first tree depth is 2, the normalized value is 2 / 5 = 0.4; if the preset maximum weight distribution entropy is 2, and the first weight distribution entropy is 1.58, the normalized value is 1.58 / 2 = 0.79. Then, according to the attention degree of the search task to the structural complexity and the weight uniformity, weights are set for the two normalized indexes, the tree depth weight is 0.4, and the weight distribution entropy weight is 0.6. The weighted sum is calculated by the index value x weight, and the weighted sum is the first search difficulty coefficient. Then, according to the preset difficulty classification threshold, 0-0.3 is simple, 0.3-0.7 is medium, and 0.7-1 is complex, the interval to which the first search difficulty coefficient belongs is determined, and the corresponding first search task level is output.
[0109] Finally, according to the first search task level, a search unit matching mechanism is used to locate the first distributed search unit. A matching mapping relationship between the search task level and the performance of the distributed search unit is established in advance, for example, the simple level matches the basic performance search unit, the medium level matches the standard performance search unit, and the complex level matches the high performance search unit, and the computing power of the search unit is also improved. Then, according to the output first search task level, the search unit meeting the performance requirement is found, and the network address, resource identifier and other information of the search unit are determined, and the location of the first distributed search unit is completed.
[0110] By combining tree traversal, information entropy calculation, and weighted normalization, the distributed search unit is accurately matched and adapted according to the complexity of the hierarchical search feature tree, and the effect of efficient execution of the search task is ensured.
[0111] Further, the method provided in the embodiment of the application comprises the following steps A600:
[0112] A610: The first distributed search unit initiates an initial query based on the top node feature word of the first hierarchical search feature tree, and outputs a first layer search result.
[0113] A620: The first distributed search unit dynamically appends the second layer node feature word of the first hierarchical search feature tree based on the first layer search result to perform a secondary query, and outputs a second layer search result.
[0114] A630: The first distributed search unit dynamically appends the third layer node feature word of the first hierarchical search feature tree based on the second layer search result to perform a tertiary query.
[0115] A640: Iteratively superimpose the feature words condition by analogy traversing the first hierarchical retrieval feature tree until output the first global retrieval result set.
[0116] A650: With the retrieval permission policy as the result desensitization constraint, remove the unauthorized field desensitization of the first global retrieval result set, and output the first refined result set.
[0117] In one embodiment, first, the first distributed retrieval unit first locates the top node feature word of the first hierarchical retrieval feature tree, which is the core entry of the tree, determined by the top node of the first hierarchical retrieval feature tree constructed in the foregoing steps, for example, the core data category of the communication data represents the retrieval. Subsequently, the retrieval unit calls the communication data retrieval engine to match all original data containing the feature word in the preset communication data repository with the top node feature word as the unique query condition, and filters out the data set that meets the core category requirement, which is the first layer retrieval result.
[0118] Then, after obtaining the first layer retrieval result, the first distributed retrieval unit no longer queries the full amount of data, but performs condition superposition based on the existing result set, extracts the second layer node feature word from the first hierarchical retrieval feature tree, which is the semantic associated child node of the top node, for example, voice communication data, further limits the data type, and dynamically appends it to the initial query rule as a new query condition. Then, the retrieval unit performs secondary screening in the first layer retrieval result, and only retains the data set containing the top node feature word + second layer node feature word, and outputs the second layer retrieval result.
[0119] Then, continuing the logic of the above-mentioned secondary query step, the first distributed retrieval unit extracts the third layer node feature word from the first hierarchical retrieval feature tree, which is the subdivided semantic node of the second layer node, for example, the near 7-day voice communication data, increases the time dimension constraint, and again dynamically appends the feature word as a query condition based on the second layer retrieval result. Subsequently, the retrieval unit performs three times of screening in the second layer retrieval result, removes the data that does not meet the combination condition of time + type + core category, and retains the data set that meets all the superimposed conditions.
[0120] Afterwards, the first distributed retrieval unit searches the hierarchical structure of the first hierarchical retrieval feature tree according to the first level, for example, the tree contains 3 layers of nodes, then traverses to the 3rd layer, and continues to execute the subsequent query operation: every time a layer is traversed, the node feature word of the layer is extracted from the tree, dynamically appended to the existing query condition combination, and then the filtering is performed in the current retrieval result set. This iterative superposition process continues until all valid levels of the tree are traversed, that is, there is no more next layer node feature word to append, and finally the data set obtained is the first global retrieval result set, which is the complete data set meeting the combination condition of all levels of the first hierarchical retrieval feature tree. The semantic retrieval condition coverage from the core to the subdivision is realized, and all the filtering is iteratively completed in the same result set, avoiding the complexity of multi-source result integration.
[0121] Finally, the first distributed retrieval unit calls the associated retrieval permission policy based on the user ID binding in step A140, for example, the ordinary user has no access right to the sensitive fields such as user ID number and communication content details, which are taken as result desensitization constraint conditions to perform field-level desensitization processing on the first global retrieval result set. The character replacement is performed on the unauthorized field, the sensitive information is masked with corresponding symbols, or the field column exceeding the permission is directly removed, and all unauthorized fields in the result set that do not meet the permission requirement are deleted or hidden.
[0122] Finally, the first refined result set is output, which fully meets the iterative filtering conditions of the first hierarchical retrieval feature tree, satisfies the security access requirements of the user permission, and is the result of dynamic condition superposition and iterative narrowing in a single hierarchical tree, rather than the integration of multi-round independent query results, ensuring the unity of data accuracy and compliance.
[0123] In summary, the communication data intelligent retrieval method based on semantic analysis provided by the embodiments of the present application has the following technical effects:
[0124] The present application binds the associated retrieval permission policy based on the user ID after receiving the original query text submitted by the user, loads the user portrait according to the user ID to perform semantic completion on the original query text to obtain a plurality of semantically equivalent completion sentences, extracts the feature words of the sentences and constructs a plurality of hierarchical retrieval feature trees in combination with the domain knowledge graph, positions a plurality of distributed retrieval units according to the node weight features of the trees, drives the units to perform condition dynamic superposition retrieval from top to bottom on the feature trees to output a plurality of refined result sets, and finally calls the result set metadata timestamp in descending order to arrange and integrate, thereby accurately outputting the communication data time sequence retrieval sequence, making the communication data intelligent retrieval result more accurate and meeting the user permission and individual needs, and achieving the technical effects of accurate semantic understanding, ordered and efficient retrieval of communication data, and accurate positioning of communication data meeting the user's intention.
[0125] Embodiment two, as Figure 2As shown, based on the same inventive concept as the preceding embodiment one, the application embodiment provides a communication data intelligent retrieval platform based on semantic analysis, which comprises:
[0126] The association retrieval permission policy execution module 1 is used to trigger the association retrieval permission policy binding based on the user ID after receiving the original query text submitted by the user.
[0127] The semantic equivalent completion sentence acquisition module 2 is used to load the user portrait according to the user ID to perform semantic completion of the original query text, and obtain a plurality of semantic equivalent completion sentences.
[0128] The retrieval feature word acquisition module 3 is used to obtain a plurality of groups of retrieval feature words by performing feature word extraction on the plurality of semantic equivalent completion sentences.
[0129] The hierarchical retrieval feature tree construction module 4 is used to perform semantic association divergence on the plurality of groups of retrieval feature words based on the domain knowledge graph to construct a plurality of hierarchical retrieval feature trees.
[0130] The distributed retrieval unit positioning module 5 is used to position a plurality of distributed retrieval units according to the node weight features of the plurality of hierarchical retrieval feature trees for hierarchical limited retrieval task matching.
[0131] The refined result set acquisition module 6 is used to drive the plurality of distributed retrieval units to perform top-down conditional dynamic superposition retrieval on the plurality of hierarchical retrieval feature trees with the retrieval permission policy as a constraint, and output a plurality of refined result sets.
[0132] The time-sequenced retrieval sequence acquisition module 7 is used to call the metadata time stamps of the plurality of refined result sets for descending order arrangement and integration, and output a communication data time-sequenced retrieval sequence.
[0133] Further, the association retrieval permission policy execution module 1 is used to perform the following steps:
[0134] Extract real-time request attributes and the user ID from the request context of the original query text, wherein the real-time request attributes include user identity / role and operating environment status; associate the user ID with the security audit database to retrieve historical unauthorized operation records and construct an operation risk tag; match the initial search permission according to the user identity / role, and then dynamically downgrade the initial search permission based on the operation risk tag, outputting the corrected search permission; perform real-time risk control to adapt the corrected search permission according to the operating environment status, and output the associated search permission strategy.
[0135] Furthermore, the semantic equivalence completion statement acquisition module 2 is used to perform the following steps:
[0136] Using the user ID as the query key, the user profile database is queried to obtain real-time user profile features; the original query text and real-time user profile features are concatenated into a structured input, and multi-branch semantic completion is performed through a semantic completion model to generate multiple semantically equivalent completion statements corresponding to multiple personalized search intent vectors.
[0137] Furthermore, the semantic equivalence completion statement acquisition module 2 is used to perform the following steps:
[0138] The real-time user profile features are converted into multi-dimensional dynamic feature vectors, wherein the multi-dimensional dynamic feature vectors have dynamic detached properties; after word segmentation from the original query text, intent anchors are inserted to generate a query intent sequence; based on the query intent sequence, the multi-dimensional dynamic feature vectors are updated by random perturbation of vector weights based on probability distribution, and the multiple personalized search intent vectors are constructed by concatenating the vectors; the original query text and the multiple personalized search intent vectors are used as structured inputs, and multi-branch semantic completion is performed through the semantic completion model to generate the multiple semantically equivalent completed sentences.
[0139] Furthermore, the semantic equivalence completion statement acquisition module 2 is used to perform the following steps:
[0140] A first intent identifier is constructed based on a first personalized retrieval intent vector, and a first structured input matrix is generated by concatenating the first intent identifier with the original query text. The semantic completion model is driven to perform bundle-width adjustable search decoding on the first structured input matrix, and output a first semantic equivalence candidate set. Time entities and core actions are extracted from the original query text, and core entity conservation rules are constructed. The semantic similarity between the first semantic equivalence candidate set and the original query text is calculated using the core entity conservation rules as constraints, so as to filter and obtain the first semantic equivalence completion statement.
[0141] Further, the hierarchical retrieval feature tree construction module 4 is configured to perform the following steps:
[0142] By performing entity linking on the first set of retrieval feature words and the nodes of the domain knowledge graph, a raw semantic layer retrieval feature is constructed. The raw semantic layer retrieval feature is used as a root node, and a semantic correlation divergence of a preset extension level is performed to obtain a first level tree node set. Based on the first level tree node set, a first initial hierarchical tree is constructed. A first historical recurrence frequency set of the first level tree node set is locally called, and a first feature word initial weight set is calculated based on the first historical recurrence frequency set. After mapping and loading the first feature word initial weight set to the first initial hierarchical tree, a weight threshold pruning is performed based on a preset hierarchical attenuation rule to construct a first hierarchical retrieval feature tree.
[0143] Further, the refined result set acquisition module 6 is configured to perform the following steps:
[0144] The first distributed retrieval unit initiates an initial query based on the top node feature word of the first hierarchical retrieval feature tree and outputs a first layer retrieval result. The first distributed retrieval unit dynamically appends the second layer node feature word of the first hierarchical retrieval feature tree to perform a secondary query based on the first layer retrieval result and outputs a second layer retrieval result. The first distributed retrieval unit dynamically appends the third layer node feature word of the first hierarchical retrieval feature tree to perform a tertiary query based on the second layer retrieval result. The feature word condition iterative superposition is iterated through the first hierarchical retrieval feature tree until a first global retrieval result set is output. The first global retrieval result set is subjected to unauthorized field desensitization removal based on the retrieval permission policy as a result desensitization constraint, and a first refined result set is output.
[0145] Further, the distributed retrieval unit positioning module 5 is configured to perform the following steps:
[0146] The first node weight set of the first hierarchical retrieval feature tree is extracted. The first tree depth and the first weight distribution entropy are calculated based on the first node weight set. The first tree depth and the first weight distribution entropy are normalized by weighting, and a first retrieval difficulty coefficient is output. The first retrieval task level is output according to the task classification based on the first retrieval difficulty coefficient. The first distributed retrieval unit is matched and positioned according to the first retrieval task level.
[0147] Further, the semantic equivalent completion sentence acquisition module 2 is configured to perform the following steps:
[0148] The real-time user portrait feature includes historical retrieval features, business preference features, and semantic habit features.
[0149] The communication data intelligent retrieval platform based on semantic analysis provided by the embodiments of the present application can execute the communication data intelligent retrieval method based on semantic analysis provided by any of the embodiments of the present application, and has the function modules and beneficial effects corresponding to the execution method.
[0150] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or server, and the various units and modules are only divided according to the functional logic, and are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual differentiation, and do not limit the protection scope of the present application.
[0151] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application. In some cases, the actions or steps described in the present application can be executed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.
Claims
1. A method for intelligent retrieval of communication data based on semantic parsing, characterized in that, The method comprises: After receiving the original query text submitted by the user, triggering the association retrieval permission policy binding based on the user ID; According to the user ID, load the user portrait to complete the semantic completion of the original query text, and obtain a plurality of semantically equivalent completion sentences; Through feature word extraction of the plurality of semantically equivalent completion sentences, a plurality of groups of retrieval feature words are obtained; Based on the domain knowledge graph, the semantic association divergence of the plurality of groups of retrieval feature words is performed to construct a plurality of hierarchical retrieval feature trees; According to the node weight features of the plurality of hierarchical retrieval feature trees, hierarchical limited retrieval task matching is performed to locate a plurality of distributed retrieval units; With the retrieval permission policy as a constraint, the plurality of distributed retrieval units are driven to perform a top-down conditional dynamic superposition retrieval of the plurality of hierarchical retrieval feature trees, and a plurality of refined result sets are output; The metadata time stamps of the plurality of refined result sets are called to perform descending order arrangement and integration, and a communication data time series retrieval sequence is output; According to the user ID, load the user portrait to complete the semantic completion of the original query text, and obtain a plurality of semantically equivalent completion sentences, the method comprising: Taking the user ID as a query key, querying the user portrait database to retrieve real-time user portrait features; The original query text and real-time user portrait features are spliced into structured inputs, and multi-branch semantic completion is performed via a semantic completion model to generate the plurality of semantically equivalent completion sentences corresponding to a plurality of personalized retrieval intent vectors; The original query text and real-time user portrait features are spliced into structured inputs, and multi-branch semantic completion is performed via a semantic completion model to generate the plurality of semantically equivalent completion sentences corresponding to a plurality of personalized retrieval intent vectors, the method comprising: Convert the real-time user portrait features into a multi-dimensional dynamic feature vector, wherein the multi-dimensional dynamic feature vector has a dynamic free attribute; From the original query text, extract the words after the segmentation and insert the intent anchor point symbols to generate a query intent sequence; According to the query intent sequence, after performing a probability distribution-based vector weight random disturbance update on the multi-dimensional dynamic feature vector, the plurality of personalized retrieval intent vectors are constructed by splicing vectors; The original query text and the plurality of personalized retrieval intent vectors are taken as structured inputs, and multi-branch semantic completion is performed via the semantic completion model to generate the plurality of semantically equivalent completion sentences; The original query text and the plurality of personalized retrieval intent vectors are taken as structured inputs, and multi-branch semantic completion is performed via the semantic completion model to generate the plurality of semantically equivalent completion sentences, the method comprising: Based on the first personalized retrieval intent vector, a first intent identifier is constructed, and a first structured input matrix is generated based on the splicing of the first intent identifier and the original query text; Drive the semantic completion model to perform search decoding with adjustable beam width on the first structured input matrix to output a first semantically equivalent candidate set; From the original query text, extract time entities and core actions to construct a core entity conservation rule; Perform semantic similarity calculation of the first semantic equivalent candidate set and the original query text, constrained by the core entity conservation rule, to obtain a first semantic equivalent completion sentence through screening; Based on the domain knowledge graph, perform semantic association divergence of the multiple sets of retrieval feature words to construct multiple hierarchical retrieval feature trees, the method comprising: By performing entity linking of the first set of retrieval feature words with the nodes of the domain knowledge graph, an original semantic layer retrieval feature is constructed; Taking the original semantic layer retrieval feature as a root node, perform semantic association divergence of a preset extension level to obtain a first level tree node set; Based on the first level tree node set, a first initial hierarchical tree is constructed; Locally call a first historical recurrence frequency set of the first level tree node set, and calculate a first feature word initial weight set based on the first historical recurrence frequency set; After mapping and loading the first feature word initial weight set to the first initial hierarchical tree, perform weight threshold pruning based on a preset hierarchical attenuation rule to construct a first hierarchical retrieval feature tree; According to the node weight features of the multiple hierarchical retrieval feature trees, perform hierarchical limited retrieval task matching to locate multiple distributed retrieval units, the method comprising: Extract a first node weight set of the first hierarchical retrieval feature tree; Calculate a first tree depth and a first weight distribution entropy based on the first node weight set; After weighted normalization of the first tree depth and the first weight distribution entropy, output a first retrieval difficulty coefficient, and then perform task classification according to the first retrieval difficulty coefficient to output a first retrieval task level; According to the first retrieval task level, locate a first distributed retrieval unit.
2. The semantic parsing based intelligent retrieval of communication data method of claim 1, wherein, After receiving the original query text submitted by the user, trigger the associated retrieval permission policy binding based on the user ID, the method comprising: Extract real-time request attributes and the user ID from the request context of the original query text, wherein the real-time request attributes include user identity roles and operation environment states; Call the historical unauthorized operation records of the user ID associated security audit library to construct an operation risk label; After matching the initial retrieval permission according to the user identity role, perform dynamic permission downgrade correction of the initial retrieval permission based on the operation risk label to output a corrected retrieval permission; According to the operation environment state, perform real-time risk control adaptation on the corrected retrieval permission to output the associated retrieval permission policy. 3.The semantic parsing based intelligent retrieval method of communication data according to claim 1, wherein, Constrained by the retrieval permission policy, drive the multiple distributed retrieval units to perform top-down conditional dynamic superposition retrieval of the multiple hierarchical retrieval feature trees to output multiple refined result sets, the method comprising: The first distributed retrieval unit initiates an initial query based on the top-level node feature word of the first hierarchical retrieval feature tree to output a first layer retrieval result; The first distributed retrieval unit dynamically appends the second layer node feature word of the first hierarchical retrieval feature tree to perform a second query based on the first layer retrieval result to output a second layer retrieval result; The first distributed retrieval unit dynamically appends the third layer node feature word of the first hierarchical retrieval feature tree to perform a third query based on the second layer retrieval result; Analogous traversal of the first hierarchical search feature tree is performed for iterative superposition of feature word conditions until a first global search result set is output. The search authority policy is used as a result desensitization constraint to remove unauthorized field desensitization from the first global search result set, and a first refined result set is output.
4. The semantic analysis based intelligent retrieval of communication data method as claimed in claim 1, wherein, The real-time user portrait features include historical search features, business preference features, and semantic habit features.
5. The intelligent retrieval platform for communication data based on semantic parsing, characterized in that, The platform for implementing the semantic analysis-based intelligent search method of communication data according to any one of claims 1-4 comprises: An associated search authority policy execution module is configured to trigger an associated search authority policy binding based on a user ID after receiving an original query text submitted by a user; A semantic equivalent completion sentence acquisition module is configured to load a user portrait based on the user ID to perform semantic completion of the original query text, thereby obtaining a plurality of semantic equivalent completion sentences; A search feature word acquisition module is configured to perform feature word extraction on the plurality of semantic equivalent completion sentences, thereby obtaining a plurality of groups of search feature words; A hierarchical search feature tree construction module is configured to perform semantic association divergence on the plurality of groups of search feature words based on a domain knowledge graph, thereby constructing a plurality of hierarchical search feature trees; A distributed search unit positioning module is configured to perform hierarchical limited search task matching based on node weight features of the plurality of hierarchical search feature trees, thereby positioning a plurality of distributed search units; A refined result set acquisition module is configured to use the search authority policy as a constraint to drive the plurality of distributed search units to perform top-down conditional dynamic superposition search on the plurality of hierarchical search feature trees, thereby outputting a plurality of refined result sets; A time-sequenced search sequence acquisition module is configured to call metadata timestamps of the plurality of refined result sets to perform descending order arrangement and integration, thereby outputting a time-sequenced search sequence of communication data.
Citation Information
Patent Citations
Sorting model establishing method and device and query automatic completion method and device
CN112528157A
Privacy enhanced intelligent search method and system based on multi-round iteration
CN120596652A