Business data analysis method and device of mobile communication network based on big data
By converting the format of user feedback information and fusing features, combined with semantic optimization and knowledge graphs, accurate response information is generated, which solves the problem of inaccurate intent extraction when users provide emotional feedback in existing technologies, and improves the efficiency and readability of user feedback processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-10
AI Technical Summary
Existing value-added service data analysis models cannot accurately extract user intent when users provide emotional feedback, leading to multiple ineffective communications, reduced after-sales efficiency, and a poor user experience.
By acquiring user feedback and service category information, performing format conversion and feature fusion, extracting keywords and sentiment keywords, and combining semantic optimization and knowledge graphs, accurate target response information is generated.
It improves the server's accuracy in understanding the user's true intent, reduces redundancy in the target response text, and enhances the efficiency and readability of user feedback processing.
Smart Images

Figure CN121833883A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to a mobile communication network service data analysis method and device based on big data. BACKGROUND
[0002] Mobile communication network service data can be divided into two categories: core communication data and value-added service data. With the evolution of network generations (2G→5G), the data type and scale have increased exponentially. As of 2024, China's mobile internet monthly traffic (DOU) reached 18.18GB / household / month, with 5G traffic accounting for nearly 60%, with an annual growth rate of over 40%.
[0003] Value-added service data analysis can further interact with users by obtaining user feedback on products and questions during product use, thereby completing mobile communication network-related business after-sales service, which is beneficial to improving current transactions and improving user satisfaction. Existing value-added service data analysis mostly uses data analysis models constructed by artificial intelligence technology. The data analysis model extracts natural speech from the collected value-added service data through natural language processing methods, and then processes related after-sales problems through simple question and answer templates. However, when users provide feedback with emotions, such as sarcasm when users are angry, the data analysis model cannot accurately extract the user's accurate intent through natural language processing methods, which may result in multiple ineffective communications, reducing after-sales efficiency and affecting user experience. SUMMARY
[0004] To address the deficiencies of the prior art, the present application provides a mobile communication network service data analysis method and device based on big data to solve the above technical problems.
[0005] In a first aspect, a mobile communication network service data analysis method based on big data is provided, comprising: Obtaining user feedback information and service category information; Converting the user feedback information to obtain first text information; Through feature fusion and semantic optimization of the first text information and service category information, target problem text information is obtained; Processing the target problem text information to obtain target reply text information; Calling a preset response template to fill in the corresponding areas of the target problem text information and the target reply text information to obtain target response information; Performing format adaptation processing on the target response information and outputting it to the user to obtain the user's response result.
[0006] Further, the user feedback information is format conversion processed to obtain first text information, including: determining the type of the user feedback information; if the user feedback information is of text type, decoding the character encoding in the user feedback information to obtain first text information; if the user feedback information is of audio type, performing speech recognition on the user feedback information to obtain first text information.
[0007] Further, the first text information and service category information are feature fused and semantically optimized to obtain target problem text information, including: performing keyword extraction processing on the first text information to obtain a target keyword set and an emotional keyword set; based on the target keyword set, the emotional keyword set and the service category information, obtaining user state information through feature fusion and emotion judgment; correcting the initial semantics of the first text information according to the user state information to obtain target semantic information corresponding to the first text information; according to the target semantic information, matching corresponding text content in a pre-set problem text database to obtain target problem text information.
[0008] Further, the first text information is keyword extraction processed to obtain a target keyword set and an emotional keyword set, including: performing word segmentation processing on the first text information to obtain a plurality of keywords; analyzing and processing the dependency relationship between the plurality of keywords to obtain a dependency relationship set; based on a target dependency relationship in the dependency relationship set, determining a target keyword set and an emotional keyword set from the plurality of keywords.
[0009] Further, based on the target keyword set, the emotional keyword set and the service category information, user state information is obtained through feature fusion and emotion judgment, including: based on the target keyword set and the emotional keyword set, evaluation label information is obtained through text vector fusion and weighted summation processing; feature vectors are extracted from the evaluation label information and the service category information respectively to obtain a first feature vector and a second feature vector; the first feature vector and the second feature vector are feature fused to obtain a user state information feature value; the user emotion polarity is determined according to the size of the user state information feature value to obtain user state information.
[0010] Further, the initial semantics of the first text information is corrected according to the user state information, to obtain target semantic information corresponding to the first text information, including: obtaining first semantic information of the first text information; extracting feature information of the first semantic information and the user state information respectively, to obtain first feature information and second feature information; based on the first feature information and the second feature information and combining a text feature vector of the first text information, calculating to obtain a target deflection value; deflecting vector information of the first semantic information according to the target deflection value, to obtain vector information of target semantic information; decoding the vector information of the target semantic information, to obtain target semantic information.
[0011] Further, the target problem text information is processed to obtain target reply text information, including: obtaining a knowledge graph corresponding to the target problem text information; extracting a reply keyword group based on the knowledge graph, to obtain a reply text information set; performing character processing on the reply text information in the reply text information set, to obtain processed reply text information; splicing the processed reply text information, to obtain target reply text information.
[0012] Further, the reply keyword group is extracted based on the knowledge graph, to obtain a reply text information set, including: extracting the target problem text information, to obtain problem keywords and search keywords; obtaining a corresponding knowledge graph according to the problem keywords and the search keywords; performing connector embedding processing on character groups in the knowledge graph and decoding, to obtain a reply text information set.
[0013] Further, the processed reply text information is spliced to obtain target reply text information, including: obtaining information lengths of each reply text information in the reply text information set; calculating the information lengths of each reply text information, to obtain an average information length; comparing the information lengths of each reply text information with the average information length, to obtain a semantic score screening; splicing the reply text information based on the semantic score screening, to obtain target reply text information.
[0014] In a second aspect, a service data analysis device for a mobile communication network based on big data is provided, which is based on any one of the service data analysis methods for a mobile communication network based on big data described above, and comprises: an acquisition module configured to acquire user feedback information and service category information; a format conversion processing module configured to perform format conversion processing on the user feedback information to obtain first text information; a feature fusion module configured to obtain target question text information through feature fusion and semantic optimization of the first text information and the service category information; a processing module configured to process the target question text information to obtain target reply text information; a filling processing module configured to call a preset answer template to perform filling processing on corresponding areas of the target question text information and the target reply text information to obtain target answer information; an output module configured to perform format adaptation processing on the target answer information and output to a user to obtain a response result of the user.
[0015] The application adopting the above technical solution has the following advantages: 1. The server of the application obtains a target deflection value between feature information of first text information corresponding to user feedback information and feature information corresponding to the user feedback information, and deflects first semantic information corresponding to the first text information according to the target deflection value to obtain target semantic information corresponding to the first text information, so that the server can accurately understand the real intention of the user and improve the accuracy of the server in obtaining target question text information corresponding to the first text information.
[0016] 2. The server of the application deletes redundant characters in a third character group to reduce redundancy in target reply text information generated by the server and provide readability of the target reply text information. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the application, the drawings required in the specific embodiments will be briefly introduced below. In all the drawings, the elements or parts are not necessarily drawn according to the actual proportions.
[0018] Fig. 1 A flowchart of the service data analysis method for a mobile communication network based on big data of the application; Fig. 2 A flowchart of the service data analysis device for a mobile communication network based on big data of the application. DETAILED DESCRIPTION
[0019] The embodiments of the technical solutions of the present application will be described in detail below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and therefore only serve as examples, but cannot limit the protection scope of the present application.
[0020] As shown in Figs. 1-2 The service data analysis method of the mobile communication network based on big data of the present application comprises: Step S01, obtaining user feedback information and service category information; Step S02, performing format conversion processing on the user feedback information to obtain first text information; Step S03, obtaining target problem text information through feature fusion and semantic optimization of the first text information and the service category information; Step S04, processing the target problem text information to obtain target reply text information; Step S05, calling a preset response template to perform filling processing on the corresponding areas of the target problem text information and the target reply text information to obtain target response information; Step S06, performing format adaptation processing on the target response information and outputting it to the user to obtain the response result of the user.
[0021] Specifically, the server obtains the user feedback information and the service category information.
[0022] The method for the server to obtain the user feedback information and the service category information can be that the server obtains the user feedback information and the service category information through the background database of the sales software. The user feedback information is the feedback information of the user on the basic service. The user feedback information can be text information or audio information. The basic service includes the sales of three types of digital products (3C), broadband service, and call service, etc. The service category information can be a complaint category, a consultation category, etc. The background database is used to store the background data in the basic service after-sales service system. The basic service after-sales service system includes a feedback information acquisition module, an artificial intelligence module, a response module, etc.
[0023] In this embodiment, the format conversion processing is performed on the user feedback information to obtain the first text information, which comprises: determining the type of the user feedback information; If the user feedback information is of a text type, the character encoding in the user feedback information is decoded to obtain the first text information; If the user feedback information is of an audio type, the user feedback information is subjected to voice recognition to obtain the first text information.
[0024] Specifically, the server obtains the first text information corresponding to the user feedback information.
[0025] The server may obtain the first text information corresponding to user feedback information through the following methods: The server determines the category of the user feedback information; if the user feedback information is text information, the server obtains the encoding information corresponding to all characters in the user feedback information; the server decodes the encoding information corresponding to all characters in the user feedback information to obtain the first text information; if the user feedback information is audio information, the server performs speech recognition on the user feedback information to obtain the first text information corresponding to the user feedback. The method used by the server to perform speech recognition on the user feedback information can be conventional speech recognition techniques, such as natural language processing, hidden Markov models, etc., which are only used as examples here and do not constitute a limitation on this solution.
[0026] In this embodiment, the target question text information is obtained through feature fusion and semantic optimization of the first text information and service category information, including: Keyword extraction is performed on the first text information to obtain a target keyword set and a sentiment keyword set; Based on the target keyword set, the sentiment keyword set, and service category information, user status information is obtained through feature fusion and sentiment judgment. The initial semantics of the first text information are corrected based on the user status information to obtain the target semantic information corresponding to the first text information. Based on the target semantic information, the corresponding text content is matched in the preset question text database to obtain the target question text information.
[0027] Specifically, the server obtains the target question text information based on the first text information and service category information.
[0028] Existing methods for obtaining question text information typically involve the server acquiring semantic information corresponding to first text information, and then matching that semantic information with the target question text information in a database. However, when a user provides feedback using sarcasm, the semantic information obtained through this method is the opposite of what the user intended to express. This reduces the accuracy of the semantic information obtained from the first text information, consequently reducing the accuracy of the target question text information. Therefore, the server can utilize an artificial intelligence module within the after-sales service system to perform the following steps to improve the accuracy of the semantic information obtained from the first text information.
[0029] The server can obtain the target question text information based on the first text information and service category information in the following ways: the server obtains the target keyword set and sentiment keyword set from the first text information; the server obtains the user's state information based on the target keyword set, sentiment keyword set, and service category information; the server obtains the target semantic information corresponding to the first text information based on the user state information; and the server obtains the target question text information based on the target semantic information. The user state information can be used to indicate the user's emotional state, such as calm or angry. By obtaining the semantic information of the first text information through the user state information, the server can ensure that the semantic information corresponding to the obtained first text is consistent with the user's intended meaning, thus better understanding the user's true intent and improving the accuracy of the obtained question text information.
[0030] In this embodiment, keyword extraction processing is performed on the first text information to obtain a target keyword set and a sentiment keyword set, including: The first text information is segmented into words to obtain multiple keywords; The dependency relationships between multiple keywords are analyzed and processed to obtain a dependency relationship set; Based on the target dependency relationships in the dependency relationship set, the target keyword set and the sentiment keyword set are determined from multiple keywords.
[0031] Specifically, the server obtains the target keyword set and sentiment keyword set from the first text information.
[0032] The server can obtain the target keyword set from the first text information using the following method: the server performs word segmentation on the first text information to obtain multiple keywords; the server obtains the dependency relationships between the keywords to obtain a dependency relationship set; and the server determines the target keyword set and the sentiment keyword set from the multiple keywords based on the dependency relationship set. The target keywords in the target keyword set correspond one-to-one with the sentiment keywords in the sentiment keyword set. The method used by the server to perform word segmentation on the first text information can be a conventional word segmentation method, such as HMM-based word segmentation or conditional random field-based word segmentation. This is only used as an example and is not intended to limit the specific solution.
[0033] Keywords may include one or more words; this scheme does not impose any restrictions on this. Dependency relationships may include subject-verb, verb-object, prepositional phrase, verb-complement, etc.
[0034] The server can determine the target keyword set and sentiment keyword set from multiple keywords based on the dependency relationship set as follows: The server obtains the target dependency relationship from the dependency relationship set; the server obtains the first keyword set and the second keyword set corresponding to the target dependency relationship; the server determines the target keyword set between the first keyword set and the second keyword set based on the target word category; the server obtains keywords with the word category of predicate in the dependency relationship set to obtain the sentiment keyword set. The target word category and target dependency relationship are determined by user settings or system defaults; the target word category can be any of the word categories of all keywords, including subject, object, predicate, attributive, adverbial, etc.; sentiment keywords can be keywords with the word category of predicate in the first text information, which can be used to reflect user evaluation; for example: average, good, bad, fluent, etc.
[0035] In a specific example, the server obtains the first text information: "This operator's 4G network speed is very stable, the 5G signal in the cell is weak, and the voice call quality is average." The server performs word segmentation on this text, obtaining 14 word segments: "this," "operator," "of," "4G network speed," "very," "stable," "in the cell," "of," "5G signal," "weak," "voice call," "of," "sound quality," and "average." The server then obtains the dependency relationships between these 14 word segments. The dependency relationship between "operator" and "of," "in the cell" and "of," and "voice call" and "of" is a function word structure; the dependency relationship between "operator" and "4G network speed," "in the cell" and "5G signal," and "voice call" and "sound quality" is a noun-head relationship; the dependency relationship between "4G network speed" and "stable," "5G signal" and "weak," and "sound quality" and "average" is a subject-predicate relationship; and the dependency relationship between "stable" and "very" is a subject-predicate relationship. The dependency relationship between the two is an adverbial-head relation, meaning the elements in the dependency relationship set are "function word structure," "attributive-head relation," "subject-predicate relation," and "adverbial-head relation." The target dependency relationship obtained by the server is a "subject-predicate relation." The elements in the first keyword set and the second keyword set obtained by the server based on the "subject-predicate relation" are "4G network speed," "5G signal," and "voice call," and "stable," "weak," and "average," respectively. Since the target word category obtained by the server is the subject, the server can use the first segmentation "4G network speed," "5G signal," and "voice call" as elements in the target keyword set of the first text information to obtain the target keyword set. The keywords obtained by the server in the dependency relationship set that are predicates are "stable," "weak," and "average," so the elements in the sentiment keyword set obtained by the server are "stable," "weak," and "average."
[0036] In this embodiment, user status information is obtained through feature fusion and emotion judgment based on the target keyword set, the sentiment keyword set, and service category information, including: Based on the target keyword set and the sentiment keyword set, evaluation tag information is obtained through text vector fusion and weighted summation. Feature vectors are extracted from the evaluation label information and service category information respectively to obtain the first feature vector and the second feature vector; The first and second feature vectors are fused to obtain the feature values of user state information. The user's emotional polarity is determined by the magnitude of the user's status information feature values, thus obtaining the user's status information.
[0037] Specifically, the server obtains the user's status information based on the target keyword set, the sentiment keyword set, and the service category information.
[0038] The method by which the server obtains user status information based on the target keyword set, the sentiment keyword set, and service category information can be as follows: the server determines evaluation tag information based on the target keyword set and the sentiment keyword set; the server obtains user status information based on the evaluation tag information and the service category information. For the specific implementation methods of the server determining evaluation tag information based on the target keyword set and the sentiment keyword set, and the server obtaining user status information based on the evaluation tag information and the service category information, please refer to the following steps.
[0039] The server determines the target evaluation tag information based on the target keyword set and the sentiment keyword set.
[0040] The server determines the evaluation tag information based on the target keyword set and the first text information using the following method: The server obtains the text vector corresponding to each target keyword in the target keyword set to obtain a target keyword text vector set; the server obtains the text vector corresponding to each sentiment keyword in the sentiment keyword set to obtain a sentiment keyword text vector set; the server performs embedding processing on each target keyword text vector in the target keyword text vector set to obtain a target keyword embedding vector set; the server fuses each target keyword embedding vector in the target keyword embedding vector set with its corresponding sentiment keyword text vector to obtain a text vector set for the first evaluation tag information; the server fuses the text vectors of the first evaluation tag information in the first evaluation tag information text vector set to obtain the text vector corresponding to the target evaluation tag information; the server decodes the text vectors of the target evaluation tag information to obtain the target evaluation tag information. Here, the sentiment keywords can be word segments in the first text information used to represent customer evaluations, such as: neutral, good, bad, etc. The method for decoding the text vectors of the evaluation tag information can be a conventional decoding method, which is not limited here. Evaluation tags can be categorized into two or more types based on requirements. For example, broadband service evaluation tags can be categorized into positive, negative, and neutral. A positive tag could be "fast internet speed," a negative tag could be "slow internet speed," and a neutral tag could be "average internet speed." Embedding text vectors involves embedding the vector of the image corresponding to the keyword into the keyword's text vector. This embedding process helps the system better understand the keyword and improves the accuracy of keyword retrieval. The target evaluation tag is the evaluation tag corresponding to the first text information. The server can fuse each target keyword embedding vector in the target keyword embedding vector set with its corresponding sentiment keyword text vector by simply concatenating the target keyword embedding vector with its corresponding sentiment keyword text vector. For example, if the target keyword embedding vector is a mobile phone and its corresponding image, and the sentiment keyword is "smooth," then the first evaluation tag obtained after concatenation would be "smooth mobile phone performance."
[0041] The method by which the server fuses the text vectors of the first evaluation label information in the first evaluation label information text vector set can be as follows: The server determines reference keywords in the target keyword set; the server obtains the weight information set corresponding to each target keyword and reference keyword in the target keyword set; the server performs a weighted summation of the first evaluation label information in the first evaluation label information set according to the weight information set to obtain the text vector of the evaluation label information. The reference keywords can be determined based on the subordinate relationship between target keywords in the target keyword set. For example, the target keywords "4G network" and "5G network" belong to a sub-service of the reference keyword "mobile communication network," therefore, "mobile communication network" in the target keyword set can be determined as the reference keyword. The target keywords and reference keywords have a subordinate relationship; therefore, the server can perform a weighted summation of the first evaluation label information corresponding to the target keywords in the target keyword set according to the weight information set to obtain the evaluation label information. The weight information corresponding to each target keyword and reference keyword in the target keyword set can be obtained through methods such as surveys to determine the importance of the item corresponding to each target keyword to the user; the higher the importance, the greater the weight value of the target keyword; the lower the importance, the smaller the weight value of the target keyword. The server can use conventional weighted summation methods, such as subjective weighting, Delphi method, and entropy weighting, to perform weighted summation on the keyword embedding vector and the text vector of sentiment keywords, based on the weight information corresponding to the target keyword and reference keywords.
[0042] The server can obtain the weight information set corresponding to each target keyword and reference keyword in the target keyword set by: the server obtaining the relevance between each target keyword and reference keyword; and the server obtaining the corresponding weight information set based on the relevance between each target keyword and reference keyword. Specifically, the relevance between each target keyword and reference keyword can be represented by calculating the product between the transpose of the vector representation of the reference keyword and the vector representation of each target keyword. A higher relevance between a target keyword and a reference keyword indicates a greater influence of the target keyword on the reference keyword, and thus a higher weight value. However, keywords and target keywords may exist in different dimensions. Therefore, each keyword and target keyword can be normalized before calculating their relevance to represent the weight information corresponding to each keyword and target keyword.
[0043] The server obtains the relevance between each keyword and the target keyword, as shown in formula (1): Formula (1) In formula (1), f represents the numerical value of the relevance between the keyword and the target keyword; M represents the vector matrix corresponding to the target keyword; The inverse matrix representing the vector matrix corresponding to the target keyword; This represents the vector matrix corresponding to the keyword.
[0044] The server obtains the set of weight information corresponding to the keywords and target keywords, as shown in formula (2): Formula (2) Among them, in formula (2) This represents the numerical value corresponding to the weight information of the keyword and the target keyword; represents the normalized exponential function; M represents the vector matrix corresponding to the target keyword; The inverse matrix representing the vector matrix corresponding to the target keyword; This represents the vector matrix corresponding to the keywords; The length of the vector representing the key.
[0045] In a specific example: The first text information obtained by the server is "This operator's 4G network speed is very stable, the 5G signal in the cell is weak, and the voice call quality is average." The keywords in the first text information obtained by the server are "4G network speed," "5G signal," and "voice call." Since the keywords "5G signal" and "voice call" both belong to the core business scope under mobile communication network services, the server can determine "4G network speed" as the target keyword. The sentiment keywords obtained by the server are "stable," "weak," and "average." The server performs vector fusion on the above keywords to obtain "4G network speed is stable," "5G signal is weak," and "voice call quality is average." According to user survey data, mobile communication users are most concerned about network speed stability, followed by voice call quality, and lastly 5G signal coverage. Based on this, the weight information corresponding to "4G network speed" is 0.5, the weight information corresponding to "voice call" is 0.3, and the weight information corresponding to "5G signal" is 0.2. The server performs weighted summation and decoding on the vectors based on the weight information to obtain the evaluation label information of the first text information as "mobile communication network user experience is average."
[0046] The server obtains user status information based on evaluation tag information and service category information.
[0047] The method by which the server obtains user status information based on evaluation tag information and service category information may be as follows: the server obtains the feature vector of evaluation tag information to obtain a first feature vector; the server obtains the feature vector of service category information to obtain a second feature vector; the server performs feature fusion on the first feature vector and the second feature vector to obtain the user status information feature value; the server obtains the user status information based on the magnitude of the user status information feature value.
[0048] The user status information feature value is a non-zero constant. A feature value greater than zero indicates a positive user status; a feature value less than zero indicates a negative user status; and a feature value close to zero indicates a neutral user status. A positive user status indicates happiness, a negative user status indicates anger, and a neutral user status indicates calmness. The service category information feature vector is also a non-zero constant. A feature vector greater than zero indicates a positive service category (e.g., positive reviews); a feature vector less than zero indicates a negative service category (e.g., negative reviews, complaints); and a feature vector close to zero indicates a neutral service category (e.g., neutral reviews, inquiries).
[0049] The server may fuse the first and second feature vectors to obtain the user state information feature values as follows: the server obtains the feature vector value of the first feature vector; the server obtains the feature vector value of the second feature vector; the server adds the feature vector value of the first feature vector to the feature vector value of the second feature vector to obtain the user state information feature values. Here, the feature vector value of the first feature vector is the numerical value corresponding to the first feature vector; the feature vector value of the second feature vector is the numerical value corresponding to the second feature vector.
[0050] Optionally, the server can further classify the obtained user status information by retrieving keywords categorized as adverbs from the first text information to obtain more accurate user status information. These adverb-categorized keywords can be modifying words, such as "very," "extremely," or "extremely." The number of levels for classifying the user status information can be determined based on user settings or system defaults. For example, if the user's status information is "angry," and the user has set the classification level to three, then "angry" can be divided into "moderately angry," "somewhat angry," and "extremely angry." "Moderately angry" corresponds to level one, "somewhat angry" to level two, and "extremely angry" to level three.
[0051] In a specific example: the server obtains the user's status information as "angry"; the keyword in the first text information obtained by the server that is an adverbial phrase is "extremely", such as the text "This operator's 5G signal is extremely poor", then the server can set the user's status level to "level three".
[0052] In another specific example: the server obtains the user's status information as "angry"; the server obtains the first text information as "This operator's data plan is unreasonable", and there are no keywords in the text that are adverbial phrases. In this case, the server can set the user's status level to "Level 1".
[0053] In this embodiment, the initial semantics of the first text information are corrected based on user state information to obtain the target semantic information corresponding to the first text information, including: Obtain the first semantic information of the first text information; The feature information of the first semantic information and the user state information are extracted respectively to obtain the first feature information and the second feature information; The target deflection value is calculated based on the text feature vector of the first feature information, the second feature information, and the first text information. The vector information of the first semantic information is deflected according to the target deflection value to obtain the vector information of the target semantic information. The target semantic information is obtained by decoding the vector information of the target semantic information.
[0054] Specifically, the server obtains the target semantic information corresponding to the first text information based on the user's state information.
[0055] The server may obtain the target semantic information corresponding to the first text information by: obtaining the semantic representation of the first text information to obtain the first semantic information; and adjusting the first semantic information according to the user's state information to obtain the target semantic information corresponding to the first text information. The method by which the server obtains the first semantic information can be a conventional method for obtaining semantic information, such as syntactic analysis or pragmatic analysis. This is merely an example and no specific limitation is imposed on the proposed solution.
[0056] The method by which the server adjusts the first semantic information based on user status information to obtain the target semantic information corresponding to the first text information can be as follows: the server obtains the vector information corresponding to the first semantic information; the server obtains the first feature information corresponding to the first semantic information; the server obtains the second feature information of the user status information; the server obtains a target deflection value based on the first text information, user status information, first feature information, and second feature information; the server deflects the vector information of the first semantic information according to the target deflection value to obtain the vector information of the target semantic information; the server decodes the vector information of the target semantic information to obtain the target semantic information corresponding to the first text information. Here, the first feature information corresponding to the first semantic information can be information representing the polarity of user evaluation, including positive, negative, neutral, etc. The deflection value is the tangent of the deflection degree; the deflection degree is the angle between the vector representation of the first semantic information and the vector representation of the target semantic information; the target deflection value is the deflection value between the first feature information and the second feature information.
[0057] The method by which the server obtains the target deflection value based on the first text information, user status information, first feature information, and second feature information may be as follows: the server obtains the text feature vector corresponding to the first text information to obtain the first text feature vector; the server obtains the projection coefficient of the first feature information on the second feature information; the server obtains the weight information corresponding to the user status information to obtain the first weight information; and the server obtains the target deflection value based on the first feature information, the second feature information, the first text feature vector, the first weight information, and the projection coefficient of the first feature information on the second feature information.
[0058] Formula (3) can be used to obtain the deflection value between the first feature information and the second feature information. Formula (3) is shown below: Formula (3) In formula (3) This represents the target deflection value between the first feature information and the second feature information; This represents the tangent operator in trigonometric functions; e is the natural exponent term; n represents the text feature vector corresponding to the first text information. This represents the weight information corresponding to user feature information; Indicates the second feature information; d represents the first feature information; d represents the projection coefficient of the vector representation of the first feature information onto the vector representation of the second feature information; k represents the deflection constant, which is set by the user; q is a random number randomly selected by the server.
[0059] Optionally, when the first text information is the text information corresponding to the audio information, the server can also obtain the audio feature information corresponding to the first text information. The spectral feature information, sound intensity information, and time information of user feedback information generated under different user information tags are all different. For example, when a user reports "This internet speed is terrible, it takes forever to load the page" under the "angry" user information tag, the corresponding spectral feature information and sound intensity information are higher than when the user inquires "I want to know when the 5G signal in my community will be available" under the "calm" user information tag. Furthermore, the time information of the feedback information in the angry state is lower than that in the calm state.
[0060] The spectral characteristics, sound intensity, and timing of user feedback differ depending on the user information tag used. For example, the spectral characteristics and sound intensity of user feedback under the "angry" tag are higher than those under the "calm" tag, while the timing of user feedback under the "angry" tag is lower than that under the "calm" tag.
[0061] Therefore, the server can obtain the audio feature information corresponding to the first text information and use formula (4) to obtain the target deflection value based on the audio feature information, the first text information, the user state information, the first feature information, and the second feature information. This allows the server to obtain the target deflection value based on multiple feature information corresponding to the first text information, thereby improving the accuracy of obtaining the target deflection value.
[0062] The method by which the server obtains the target deflection value based on audio feature information, first text information, user status information, first feature information, and second feature information may be as follows: the server obtains the change curve between audio feature information and time information based on the time information in the audio feature information; the server obtains the weight information corresponding to the user status information; the server obtains the projection coefficient of the first feature information on the second feature information; the server obtains the deflection value between the first feature information and the second feature information based on the change curve between audio feature information and time information, the weight information corresponding to the user status information, and the projection coefficient of the first feature information on the second feature information, so as to obtain the target deflection value.
[0063] Formula (4) is shown below: Formula (4) In formula (4), S represents the target deflection value between the first feature information and the second feature information; This represents the tangent operator in trigonometric functions; e is the natural exponent term; x is the audio feature information corresponding to the first text information, excluding time information. This represents the weight value corresponding to the user status information. The weight value corresponding to the user status information can be determined by user settings or by system default. The function representing the change between the audio feature information and the time information corresponding to the first text information; This represents the time information within the audio feature information; Represents the differential operator; Indicates the second feature information; d represents the first feature information; d represents the projection coefficient of the vector representation of the first feature information onto the vector representation of the second feature information; k represents the deflection constant, which is set by the user; q is a random number randomly selected by the server.
[0064] The server obtains the target question text information based on the target semantic information.
[0065] One method for a server to obtain target question text information based on target semantic information is by retrieving the target question text information from a question text database. This question text database is a pre-entered database in the basic business after-sales service system used to store a large amount of response text information, and it is accessible to the server at any time.
[0066] In this embodiment, the target question text information is processed to obtain the target response text information, including: Obtain the knowledge graph corresponding to the text information of the target question; The response keyword groups are extracted based on the knowledge graph to obtain the response text information set; The character processing is performed on the reply text information in the reply text information set to obtain the processed reply text information; The processed reply text information is concatenated to obtain the target reply text information.
[0067] Specifically, the server obtains the target response text information based on the target question text information.
[0068] Existing methods for obtaining target response text information typically involve retrieving matching response text from a target response text database based on the target question text. However, target text obtained through simple matching can be redundant, reducing readability. Therefore, the server can utilize the artificial intelligence module of the after-sales service system to perform the following steps to improve the readability of the obtained target response text information.
[0069] The server can obtain the target response text information based on the target question text information in the following ways: the server obtains the response text information corresponding to each question text information in the target question text information to obtain a set of response text information; the server concatenates the response text information in the response text information set to obtain the target response text information.
[0070] In this embodiment, the response keyword groups are extracted based on the knowledge graph to obtain a set of response text information, including: Extract the target question text information to obtain question keywords and search keywords; Based on the question keywords and search keywords, the corresponding knowledge graph is obtained; The character groups in the knowledge graph are processed by connector embedding and then decoded to obtain a set of response text information.
[0071] Specifically, the server retrieves the knowledge graph corresponding to each question text in the target question text information.
[0072] The server can obtain the knowledge graph corresponding to each question text in the target question text information by: the server obtaining the question keywords corresponding to each question text in the target question text information; the server obtaining the search keywords corresponding to each question text in the target question text information; and the server obtaining the knowledge graph corresponding to each question text information based on the question keywords and search keywords corresponding to each question text information.
[0073] The method for obtaining the question keywords and search keywords corresponding to each question text in the target question text information can refer to the method for obtaining target keywords and sentiment keywords in the above steps, and will not be repeated here. Question keywords can be keywords in the target question text information that are the subject of the word category, such as "5G signal". Search keywords can be keywords in the target question text information that are the predicate of the word category, such as "weak signal" and "poor coverage". The knowledge graph corresponding to each question text information can be composed of some keywords. The head entity of the knowledge graph can be composed of question keywords in the question text information, such as "5G signal", and the tail entity is the reply keywords in the historical reply text, such as "restart the phone", "contact the operator to check the base station", and "switch to 4G network". The head entity and tail entity of the knowledge graph are connected through the search keyword "weak signal", so the knowledge represented by the knowledge graph can be "when the 5G signal is weak, you can restart the phone, switch to 4G network or contact the operator to check the base station". Historical question text information includes question text information that is the same as or similar to the question text in the question text information database, and historical reply information can be reply text information corresponding to historical question text information. In a knowledge graph, the question keyword in the head entity is unique, and the tail entity corresponding to the head entity can include multiple response keywords; the order of the multiple response keywords in the tail entity is the same as the order in which the multiple response keywords appear in the historical response text information.
[0074] The server retrieves the corresponding response text information for each question text information based on the knowledge graph corresponding to each question text information.
[0075] The server can obtain the corresponding response text information for each question text information based on the knowledge graph corresponding to each question text information in the following ways: The server obtains a first character group based on the knowledge graph corresponding to each question text information; the server performs connector embedding processing on the first character group to obtain a second character group; the server obtains the confidence score between each connector in the second character group and a character in the character database; the server replaces the connectors in the second character group with the character in the character database with the character that has the highest confidence score with them to obtain a third character group; the server deletes redundant characters in the third character group to obtain a fourth character group; the server decodes the fourth character group to obtain the response information corresponding to the question text information. Here, the characters in the character group can be a word segment or a vector encoding corresponding to a character. Connectors can be replaced by any character in the character database or a blank (no character). The character database is pre-set by the user, and the server can access the character database at any time. The character database stores a large number of characters, which can be any of the following: English letters, Chinese characters, punctuation marks, and numbers. The confidence level between the connector and the characters in the character database can be obtained by calculating the confidence level between the two characters adjacent to the connector in the second character group and the characters in the character database; the confidence level between each connector in the second character group and the characters in the character database can be obtained using conventional confidence level calculation methods. The order of the characters in the first character group is the same as the order of the multiple response keywords in the tail entity of the knowledge graph.
[0076] The server can perform concatenation embedding on the first character group to obtain the second character group in the following ways: The server normalizes the characters in the first character group to obtain an intermediate character group. The server then obtains the probability distribution of connectors at each position in the intermediate character group. The server determines whether the probability distribution of connectors at each position is greater than a preset probability threshold. If the probability distribution of a connector is greater than the preset threshold, the server embeds a connector at that position. If the probability distribution of a connector is less than or equal to the preset threshold, the server does not perform any processing at that position. The server embeds a connector at each position in the intermediate character group where the probability distribution is greater than the threshold, to obtain a second character group. Each position in the intermediate character group includes the position between characters, the start position, and the end position. The server normalizes the characters in the first character group using a pre-input projection matrix. The probability distribution of connectors at each position can be obtained through analysis of historical data.
[0077] Therefore, the server can determine the target number of connectors in the intermediate character group according to the distribution probability of the connectors at each permutation position through formula (5). Formula (5) is shown as follows: Formula (5) In formula (5), represents the distribution probability of the connectors at each permutation position; represents the target number of connectors in the intermediate character group; represents the i-th permutation position of the connectors in the intermediate character group; represents the intermediate character group; represents the natural exponential term; represents the pre-input projection matrix; represents the pre-input projection matrix which is the transpose matrix of; represents the i-th character and the (i + 1)-th character in the intermediate character group.
[0078] If the server directly decodes the third character group, there may be a situation where redundant characters appear in the obtained reply text information, resulting in redundant content and reduced readability of the reply text information. For example: the reply text information obtained by the server is "Please restart your mobile device". If the word "once" in the reply text information is deleted, it will not change the semantic information corresponding to the reply text information. Therefore, the word "once" in the reply text information will cause redundant text content and reduced readability. To solve the above problem of redundant content in the reply text information, the server can delete the redundant characters in the third character group to delete the redundant characters in the reply text information and improve the readability of the reply text information. For example, the server can train a character deletion model to perform the method of deleting redundant characters through this character deletion model, and this solution is not limited thereto.
[0079] The method for the server to delete the redundant characters in the third character group can be as follows: The server obtains the accuracy rate of each character in the third character group; the server determines whether the character needs to be deleted according to the accuracy rate of each character in the third character group; if the accuracy rate corresponding to the character in the third character group is less than or equal to the preset accuracy rate threshold, the server deletes the character; if the accuracy rate corresponding to the character in the third character group is greater than the preset accuracy rate threshold, the server retains the character. Among them, the accuracy rate of each character is used to represent the influence rate of the character on the semantic information corresponding to the reply text information; if the change amount of the semantic information corresponding to the reply text information after the character is deleted is larger, the accuracy rate corresponding to the character is larger; if the change amount of the semantic information corresponding to the reply text information after the character is deleted is smaller, the accuracy rate corresponding to the character is smaller. For example: The target reply text information corresponding to the decoded third character group obtained by the server is "Please restart your mobile device"; through calculation, it is found that the accuracy rate corresponding to the encoding vector of "once" in the third character group is 0.3, and the preset accuracy rate threshold is 0.6; while the accuracy rates of the encoding vectors corresponding to "Please", "restart", "your", "mobile", and "device" are all greater than the preset threshold, and the server deletes the encoding vector corresponding to "once" to obtain the fourth character group; the target reply text information obtained after decoding is "Please restart your mobile device".
[0080] Among them, the server can use formula (6) to obtain the accuracy rate of each character in the third character group. Formula (6) is as follows: Formula (6) In formula (6) represents the accuracy rate of the i-th character in the target reply text information, which can be obtained by statistics in the historical database; represents the probability distribution that the accuracy rate taking the maximum probability is the accuracy rate of the i-th character in the third character group; represents the target reply text information; T is the model parameter pre-input by the character deletion model adopted by the server; represents the logistic regression function; is the vector projection matrix of the target reply text information, and the method for vector projection of the target reply text information can be a conventional vector projection method; represents the i-th character in the target reply text information.
[0081] In this embodiment, the processed reply text information is spliced to obtain the target reply text information, including: Obtain the information lengths of each reply text information in the reply text information set; Calculate the information lengths of each reply text information to obtain the average information length; The semantic score is used to filter responses by comparing the length of each response text with the average length. The target response text information is obtained by filtering and concatenating the response text information based on semantic scoring.
[0082] Specifically, the server concatenates the reply text information from the reply text information set to obtain the target reply text information.
[0083] The set of response text information obtained through the above steps may contain two or more responses with identical or similar semantics. If these semantically identical or similar responses are not processed, the target response text will contain complex and difficult-to-understand sentences, reducing its readability. Therefore, by concatenating semantically identical or similar responses in the set, the readability of the target response text can be improved.
[0084] The server can concatenate the reply text information in the reply text information set as follows: The server obtains the information length of each reply text information in the reply text information set; the server obtains the average information length of the reply text information set based on the information length of each reply text information; it determines whether the information length of each reply text information in the reply text information set is greater than the average information length. If it is greater than the average information length, a first-class punctuation mark is added to each reply text information with a length greater than the average information length to obtain a first reply text information set; if it is less than or equal to the average information length, a second-class punctuation mark is added to each reply text information with a length less than or equal to the average information length to obtain a second reply text information set; the server obtains the first semantic score corresponding to each second reply text information in the second reply text information set; the server randomly selects a first-class semantic score from the second reply text information set. The server first obtains a set of two response texts to obtain a response text to be processed. Then, it obtains a response text to be determined based on the set of the response text to be processed and the second set of response texts. Next, the server obtains a second semantic score for the response text to be determined. Finally, the server determines whether the second semantic score of the response text to be determined is greater than the first semantic score corresponding to the response text to be processed. If the semantic score is greater than the first semantic score, the response text to be determined is identified as the second target response text. If the semantic score is less than or equal to the first semantic score, the response text to be processed is identified as the second target response text. The server then arranges the second target response text and the first response text according to their original order to obtain the target response text.
[0085] The system uses two punctuation marks: a first category includes question marks, periods, exclamation marks, and ellipses; a second category includes commas, pauses, semicolons, colons, single quotes, and double quotes. Both categories are added to the end of the response text. The first and second semantic scores can be obtained through a pre-input semantic scoring model. The server obtains the response text to be determined based on the set of response text to be processed and the set of second response text. This can be achieved by placing the second response text from the set of second response text at the beginning of the response text to be processed. The response text to be determined includes the response text to be processed and at least one second response text. The semantic score corresponding to the response text reflects the semantic completeness of the response text. Higher semantic completeness results in a higher semantic score, and vice versa. After determining the response text to be processed, the second response text in the set of second response text no longer includes the response text to be processed. The order of the second target response text information is the same as the order of the second response text information at the beginning of the second target response text information.
[0086] By determining whether the length of each reply text in the reply text set is greater than the average length, corresponding punctuation marks are added to the reply text based on the determination result to avoid long and difficult sentences that are hard to understand. By determining whether the second semantic score of the reply text to be determined is greater than the first semantic score of the reply text to be processed, it is ensured that the semantics of the concatenated reply text do not change, making the obtained target reply text more accurate.
[0087] The server determines the target response information based on the target question text information and the target response text information.
[0088] One method for the server to determine the target response information based on the target question text information and the target response text information is as follows: The server fills the target question text information and the target response text information into the corresponding positions according to the pre-input response template to obtain the target response information.
[0089] The server uses the target response information to respond to the user.
[0090] The server can either send the target response information in text form to the dialog box of the after-sales service information platform, or it can convert the target response information into audio information according to the pre-input tone to obtain the target audio information, and then send the target audio information to the dialog box of the after-sales service information platform to complete the response to the user.
[0091] In other embodiments, a service data analysis apparatus for a big data-based mobile communication network is provided, and a service data analysis method for a big data-based mobile communication network based on any of the preceding embodiments is provided, comprising: The acquisition module is configured to acquire user feedback information and service category information. The format conversion processing module is configured to perform format conversion processing on user feedback information to obtain the first text information; The feature fusion module is configured to obtain the target question text information by feature fusion and semantic optimization of the first text information and service category information; The processing module is configured to process the target question text information to obtain the target response text information; The fill processing module is configured to call a preset response template and fill the corresponding areas of the target question text information and the target response text information to obtain the target response information; The output module is configured to perform format adaptation processing on the target response information and output it to the user to obtain the user's response result.
[0092] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A method for analyzing service data in a mobile communication network based on big data, characterized in that, include: Obtain user feedback information and service category information; The user feedback information is processed by format conversion to obtain the first text information; By fusing features and semantically optimizing the first text information and service category information, the target question text information is obtained; The target question text information is processed to obtain the target response text information; A preset response template is invoked to fill in the corresponding areas of the target question text information and the target response text information to obtain the target response information; The target response information is processed for format adaptation and output to the user to obtain the user's response result.
2. The service data analysis method for mobile communication networks based on big data according to claim 1, characterized in that, The user feedback information is format-converted to obtain first text information, including: Determine the type of the user feedback information; If the user feedback information is text, then the character encoding in the user feedback information is decoded to obtain the first text information; If the user feedback information is in the form of audio, then speech recognition is performed on the user feedback information to obtain the first text information.
3. The service data analysis method for mobile communication networks based on big data according to claim 1, characterized in that, By fusing features and semantically optimizing the first text information and service category information, the target question text information is obtained, including: The first text information is processed by keyword extraction to obtain a target keyword set and a sentiment keyword set; Based on the target keyword set, the sentiment keyword set, and the service category information, user status information is obtained through feature fusion and sentiment judgment. The initial semantics of the first text information are corrected based on the user status information to obtain the target semantic information corresponding to the first text information; Based on the target semantic information, the corresponding text content is matched in a preset question text database to obtain the target question text information.
4. The service data analysis method for mobile communication networks based on big data according to claim 3, characterized in that, The first text information is processed by keyword extraction to obtain a target keyword set and a sentiment keyword set, including: The first text information is segmented into words to obtain multiple keywords; The dependency relationships among the multiple keywords are analyzed and processed to obtain a dependency relationship set; Based on the target dependency relationships in the dependency relationship set, a target keyword set and a sentiment keyword set are determined from multiple keywords.
5. The service data analysis method for mobile communication networks based on big data according to claim 3, characterized in that, Based on the target keyword set, the sentiment keyword set, and service category information, user state information is obtained through feature fusion and sentiment judgment, including: Based on the target keyword set and the sentiment keyword set, evaluation tag information is obtained through text vector fusion and weighted summation. Feature vectors are extracted from the evaluation label information and service category information respectively to obtain a first feature vector and a second feature vector; The first feature vector and the second feature vector are fused to obtain the user state information feature value; The user's emotional polarity is determined based on the magnitude of the user status information feature values, thus obtaining the user status information.
6. The service data analysis method for mobile communication networks based on big data according to claim 3, characterized in that, The initial semantics of the first text information are corrected based on the user state information to obtain the target semantic information corresponding to the first text information, including: Obtain the first semantic information of the first text information; The feature information of the first semantic information and the user state information are extracted respectively to obtain the first feature information and the second feature information; The target deflection value is calculated based on the first feature information, the second feature information, and the text feature vector of the first text information. The vector information of the first semantic information is deflected according to the target deflection value to obtain the vector information of the target semantic information; The target semantic information is obtained by decoding the vector information of the target semantic information.
7. The service data analysis method for mobile communication networks based on big data according to claim 1, characterized in that, The target question text information is processed to obtain the target response text information, including: Obtain the knowledge graph corresponding to the target question text information; Based on the knowledge graph, the response keyword groups are extracted to obtain the response text information set; The reply text information in the set of reply text information is processed to obtain the processed reply text information; The processed response text information is concatenated to obtain the target response text information.
8. The service data analysis method for mobile communication networks based on big data according to claim 7, characterized in that, Based on the knowledge graph, the response keyword groups are extracted to obtain a set of response text information, including: The target question text information is extracted to obtain question keywords and search keywords; Based on the question keywords and search keywords, the corresponding knowledge graph is obtained; The character groups in the knowledge graph are processed by connector embedding and then decoded to obtain a set of response text information.
9. The service data analysis method for mobile communication networks based on big data according to claim 7, characterized in that, The processed response text information is concatenated to obtain the target response text information, including: Obtain the information length of each reply text in the reply text information set; The average message length is obtained by calculating the message length of each reply text. The semantic score is used to filter responses by comparing the length of each response text with the average length. The target response text information is obtained by filtering and concatenating the response text information based on semantic scoring.
10. A service data analysis device for a mobile communication network based on big data, characterized in that, A service data analysis method for a big data-based mobile communication network according to any one of claims 1 to 9, comprising: The acquisition module is configured to acquire user feedback information and service category information. The format conversion processing module is configured to perform format conversion processing on the user feedback information to obtain first text information; The feature fusion module is configured to obtain the target question text information by feature fusion and semantic optimization of the first text information and service category information; The processing module is configured to process the target question text information to obtain the target response text information; The filling processing module is configured to call a preset response template to fill the corresponding areas of the target question text information and the target response text information to obtain the target response information; The output module is configured to perform format adaptation processing on the target response information and output it to the user to obtain the user's response result.