Power grid data pushing system, method and equipment based on multimedia interaction and medium
Through a power grid data push system based on multimedia interaction and using face recognition and voice splitting technology, the problem that existing power grid data push methods cannot accurately match personalized needs is solved, and accurate push of power grid data and improved work efficiency are achieved.
Patent Information
- Application Number
- CN202510588339.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-19
AI Technical Summary
Existing methods for pushing power grid data cannot accurately push data according to the personalized needs of personnel in different positions, resulting in information redundancy and reducing the work efficiency of power grid personnel.
A power grid data push system based on multimedia interaction is adopted, including a power grid management center, a data acquisition module, a face recognition module, a voice splitting module and a voice recognition module. Through face recognition and voice splitting technology, it can identify user intentions and accurately push relevant power grid data.
It achieves accurate identification of user intentions, avoids the push of useless data, and improves the accuracy of power grid data push and the work efficiency of power grid personnel.
Smart Images

Figure CN120670652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grid data interaction, and in particular to a power grid data push system, method, device and medium based on multimedia interaction. Background Art
[0002] With the rapid development of the power industry and the continuous expansion of power grids, the importance of grid data has become increasingly prominent. In the context of smart grid construction, grid operations generate massive amounts of data, covering a wide range of aspects, including equipment operating status, power load fluctuations, and user electricity usage behavior. This data is not only critical for stable grid operation but also serves as a vital basis for power companies to make informed decisions, optimize resource allocation, and improve service quality. To better utilize this data, grid data push technology has emerged. Its purpose is to deliver relevant data to those who need it in a timely manner, enabling them to quickly respond and make decisions.
[0003] However, current methods for pushing power grid data have significant limitations. Existing push mechanisms primarily rely on fixed trigger conditions, such as only pushing data when equipment operating parameters reach a set threshold. However, the power system involves a wide range of personnel in diverse roles, and their needs for grid data vary greatly. For power maintenance personnel, data on equipment failures and abnormal operating conditions is crucial. This data helps them promptly identify and address equipment issues, ensuring the safe and stable operation of the power grid. Power planners, on the other hand, require data such as long-term load forecasts and regional power demand growth trends to formulate scientific and rational power grid planning. The existing "one-size-fits-all" push method fails to accurately meet the individual needs of different personnel. This results in power maintenance personnel receiving a large amount of data irrelevant to equipment failures, increasing information screening costs. This makes it difficult for power planners to obtain data valuable for planning. This disconnect between data and needs creates a significant amount of information redundancy, distracting grid personnel, ultimately reducing work efficiency, and hindering the efficient operation and sustainable development of the power system. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by the present invention is: how to solve the problem that the current existing power grid data push method relies on fixed conditions to trigger and cannot accurately push according to the personalized needs of personnel in different positions, resulting in information redundancy and reduced work efficiency of power grid personnel.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: a power grid data push system based on multimedia interaction, which includes a power grid management center, a data acquisition module, a face recognition module, a voice splitting module, a voice recognition module and a data push module;
[0007] The power grid management center is used to manage each module;
[0008] The data acquisition module is used to respond to the data push request initiated by the user, collect user images and multimedia interactive voice based on the data push request, and send them to the face recognition module and the voice separation module;
[0009] The face recognition module receives the user image, performs face recognition on the user, obtains the face recognition result and sends it to the voice separation module;
[0010] The voice splitting module receives the face recognition result and the multimedia interactive voice, and if the user passes the identity authentication, the multimedia interactive voice is sequence-splitted to obtain multiple target voice pulse subsequences and sent to the voice recognition module;
[0011] The speech recognition module receives the target speech pulse subsequence, obtains the target intention recognition result and sends it to the data push module;
[0012] The data push module receives and splices the target intention recognition results to obtain a voice intention recognition result, and pushes the target power grid data corresponding to the voice intention recognition result to the user terminal.
[0013] As a preferred solution of the multimedia interactive power grid data push system described in the present invention, wherein: the multimedia interactive voice is sequence-splitting to obtain multiple target voice pulse subsequences and sending them to the voice recognition module, including:
[0014] Construct a speech state transition probability matrix based on the voiceprint information of the user's historical interactive speech;
[0015] The matrix elements in the voice state transition probability matrix represent the state transition probability of transferring from the current voiceprint category to the next voiceprint category;
[0016] Performing speech analysis on the pulse sequence of the multimedia interactive speech based on the speech state transition probability matrix to obtain a plurality of speech sequence splitting points of the multimedia interactive speech;
[0017] The pulse sequence of the multimedia interactive speech is sequence-splitting according to the plurality of speech sequence splitting points to obtain a plurality of target speech pulse subsequences.
[0018] As a preferred solution of the multimedia interactive power grid data push system described in the present invention, wherein: based on the voice state transition probability matrix, the pulse sequence of the multimedia interactive voice is subjected to voice analysis to obtain multiple voice sequence splitting points of the multimedia interactive voice, including:
[0019] Obtaining the time interval and amplitude difference of each pulse in the pulse sequence of the multimedia interactive voice;
[0020] Starting from the starting pulse of the pulse sequence of the multimedia interactive speech, calculating the transition path of the voiceprint category corresponding to each pulse based on the voice state transition probability matrix and the time interval and amplitude difference of each pulse, to obtain the voice state transition path of the pulse sequence of the multimedia interactive speech;
[0021] On the speech state transition path, if the state transition probability of the voiceprint category corresponding to the target pulse position is less than a preset probability threshold, the target pulse position is determined as the speech sequence splitting point.
[0022] As a preferred solution of the multimedia interactive power grid data push system described in the present invention, the speech recognition module receives the target speech pulse subsequence and obtains the target intention recognition result, including:
[0023] Combining multiple continuous target speech pulse subsequences to obtain a speech pulse subsequence combination;
[0024] Inputting each target speech pulse subsequence and the corresponding speech pulse subsequence combination into the intention recognition model for speech recognition, obtaining a first initial recognition result for each target speech pulse subsequence and a second initial recognition result for each speech pulse subsequence combination;
[0025] The first initial recognition result includes a first initial intention recognition result and a corresponding first intention recognition score;
[0026] The second initial recognition result includes a second initial intention recognition result and a corresponding second intention recognition score;
[0027] Filtering the second initial recognition result based on the first intention recognition score and the second intention recognition score to obtain an intention screening result;
[0028] If the intention screening result includes at least one third initial recognition result, the first initial intention recognition result and the third initial intention recognition result in the third initial recognition result are fused to obtain the target intention recognition result corresponding to each target speech pulse subsequence.
[0029] As a preferred solution of the multimedia interactive power grid data push system described in the present invention, the first initial intention recognition result and the third initial intention recognition result in the third initial recognition result are integrated, including:
[0030] Deconstructing the first initial intent recognition result and the third initial intent recognition result according to preset semantic units to obtain a first semantic unit set and a second semantic unit set respectively;
[0031] Based on the semantic relevance of each element in the first semantic unit set and the second semantic unit set, the first semantic unit set and the second semantic unit set are recombined to obtain recombined semantic units and construct an intention relationship network;
[0032] The edges of each node in the intention relationship network are calculated based on semantic similarity and co-occurrence frequency. After multiple rounds of information propagation, fused information is obtained to obtain the target intention recognition result of the target speech pulse subsequence.
[0033] As a preferred solution of the multimedia interactive power grid data push system described in the present invention, it also includes:
[0034] The intention recognition score in the third initial recognition result is greater than or equal to the intention recognition score in the first initial recognition result;
[0035] If the intention screening result is an empty set, the first initial intention recognition result in the first initial recognition result is determined as the target intention recognition result of each target speech pulse subsequence.
[0036] As a preferred solution of the multimedia interactive power grid data push system described in the present invention, wherein: each node edge of the intention relationship network is calculated based on semantic similarity and co-occurrence frequency, and fusion information is obtained through multiple rounds of information propagation to obtain the target intention recognition result of the target speech pulse subsequence, including:
[0037] Perform information propagation and feature extraction on each node of the intention relationship network to obtain a first node intention feature of the fused node information, and determine an intention feature set based on the first node intention feature;
[0038] Performing semantic expansion on each target intent feature in the intent feature set based on a preset semantic knowledge base to construct a semantic feature cluster;
[0039] Perform topic mining based on the semantic feature cluster to obtain the intended topic;
[0040] The target intention recognition result corresponding to each target speech pulse subsequence is obtained based on the probability distribution of the intention theme.
[0041] Another object of the present invention is to provide a method for pushing power grid data based on multimedia interaction.
[0042] To solve the above technical problems, the present invention provides the following technical solutions: a method for pushing power grid data based on multimedia interaction, comprising: responding to a data push request initiated by a user, and collecting user images and multimedia interactive voice based on the data push request;
[0043] Performing facial recognition on the user to obtain a facial recognition result; if the user passes the identity verification, the multimedia interactive voice is sequence-splitting to obtain multiple target voice pulse subsequences;
[0044] Obtain target intention recognition results based on the target speech pulse subsequence;
[0045] The target intention recognition results are spliced to obtain a voice intention recognition result, and the target power grid data corresponding to the voice intention recognition result is pushed to the user terminal.
[0046] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, it implements the steps of the power grid data push system based on multimedia interaction.
[0047] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the power grid data push system based on multimedia interaction are implemented.
[0048] The beneficial effects of the present invention are as follows: the voice intention recognition result can be accurately identified based on the user's multimedia interactive voice, thereby accurately pushing the target power grid data corresponding to the voice intention recognition result to the user, avoiding the push of a large amount of useless data, and improving the accuracy of power grid data push. Therefore, users can efficiently obtain the required power grid data, further improving the work efficiency of power grid personnel. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0050] Figure 1 This is a logical diagram of the system architecture of a multimedia interactive power grid data push system according to an embodiment of the present invention.
[0051] Figure 2 The figure is a schematic diagram of the overall process of a power grid data push system based on multimedia interaction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0052] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0053] Example 1, reference Figure 1 , is an embodiment of the present invention, which provides a power grid data push system based on multimedia interaction, including: a power grid management center, a data acquisition module, a face recognition module, a voice splitting module, a voice recognition module and a data push module;
[0054] The power grid management center is used to manage each module;
[0055] In an embodiment of the present invention, the power grid management middle station can be understood as a server. Therefore, the power grid management middle station continuously monitors the network port and waits for the data push request initiated by the user. When the user triggers the data push operation on his terminal device, such as a mobile phone APP, a computer client, etc., the request is sent to the power grid management middle station through the network. Therefore, after receiving the data push request, the data acquisition module responds to the data push request initiated by the user, sends an instruction to the terminal device, starts the camera and collects the user's current face image. For example, for a mobile phone APP, the APP calls the mobile phone camera, takes a high-definition front-facing photo and uploads it to the data acquisition module to obtain the user's user image.
[0056] The data acquisition module is used to respond to the data push request initiated by the user, collect the user's image and multimedia interactive voice based on the data push request, and send them to the face recognition module and the voice segmentation module;
[0057] The face recognition module is used to receive user images, perform face recognition on the user, obtain face recognition results and send them to the voice separation module;
[0058] In one possible embodiment, the facial recognition module preprocesses the user image, including adjusting image resolution, grayscaling, and removing noise, to produce a preprocessed image. Furthermore, the facial recognition module uses a preset facial recognition algorithm, such as a deep learning-based convolutional neural network algorithm, to extract facial features from the preprocessed image. This includes analyzing and extracting the contours and positional relationships of key areas such as the eyes, nose, and mouth, to generate a set of representative feature vectors. Furthermore, the facial recognition module compares the extracted feature vectors with user facial feature templates pre-stored in the power grid user database to determine a feature vector similarity value.
[0059] If it is determined that the similarity value of the feature vector with at least one of the user's facial feature templates is greater than or equal to the preset similarity threshold, the user's facial recognition result is determined to have passed identity verification. If it is determined that the similarity values of the feature vectors with the user's facial feature templates are all less than the preset similarity threshold, the user's facial recognition result is determined to have failed identity verification.
[0060] The voice splitting module is used to receive the face recognition results and multimedia interactive voice. If the user passes the identity authentication, the multimedia interactive voice is split into sequences to obtain multiple target voice pulse subsequences and send them to the voice recognition module;
[0061] In an embodiment of the present invention, the steps of obtaining multiple target speech pulse subsequences include:
[0062] Construct a speech state transition probability matrix based on the voiceprint information of the user's historical interactive speech;
[0063] The matrix elements in the voice state transition probability matrix represent the state transition probability of transferring from the current voiceprint category to the next voiceprint category;
[0064] Perform speech analysis on the pulse sequence of multimedia interactive speech based on the speech state transition probability matrix to obtain multiple speech sequence splitting points of the multimedia interactive speech;
[0065] The pulse sequence of multimedia interactive speech is sequence-splitting according to a plurality of speech sequence splitting points to obtain a plurality of target speech pulse subsequences.
[0066] In a possible embodiment, multiple speech sequence splitting points of the multimedia interactive speech may be obtained as follows:
[0067] Obtain the time interval and amplitude difference of each pulse in the pulse sequence of multimedia interactive speech;
[0068] Starting from the initial pulse of the multimedia interactive speech pulse sequence, the transition path of the voiceprint category corresponding to each pulse is calculated based on the speech state transition probability matrix and the time interval and amplitude difference of each pulse, thereby obtaining the speech state transition path of the multimedia interactive speech pulse sequence;
[0069] On the speech state transition path, if the state transition probability of the voiceprint category corresponding to the target pulse position is less than a preset probability threshold, the target pulse position is determined as the speech sequence splitting point.
[0070] The speech recognition module is used to receive the target speech pulse subsequence, obtain the target intention recognition result and send it to the data push module;
[0071] In an embodiment of the present invention, the steps of receiving a target speech pulse subsequence and obtaining a target intention recognition result include:
[0072] Combining multiple continuous target speech pulse subsequences to obtain a speech pulse subsequence combination;
[0073] In a possible embodiment, the target speech pulse subsequence is V i , the speech pulse subsequence combination can be:
[0074] {V i-2 ,V i-1 ,V i},{V i-1 ,V i ,V i+1} and {V i ,V i+1 ,V i+2}
[0075] Among them, V i-2 represents the first two segments of the target speech pulse subsequence, V i-1 represents the previous segment of the target speech pulse subsequence, V i+1 represents the next segment of the target speech pulse subsequence, V i+2 Represents the last two speech pulse subsequences of the target speech pulse subsequence.
[0076] Inputting each target speech pulse subsequence and the corresponding speech pulse subsequence combination into the intention recognition model for speech recognition, obtaining a first initial recognition result for each target speech pulse subsequence and a second initial recognition result for each speech pulse subsequence combination;
[0077] It should be noted that the speech recognition module is pre-trained with an intent recognition model, which is trained based on sample speech pulse sequences and their corresponding intent result labels.
[0078] The first initial recognition result includes a first initial intention recognition result and a corresponding first intention recognition score;
[0079] The second initial recognition result includes a second initial intention recognition result and a corresponding second intention recognition score;
[0080] Filtering the second initial recognition result based on the first intention recognition score and the second intention recognition score to obtain an intention filtering result;
[0081] If the intention screening result includes at least one third initial recognition result, the first initial intention recognition result and the third initial intention recognition result in the third initial recognition result are fused to obtain the target intention recognition result corresponding to each target speech pulse subsequence.
[0082] In an embodiment of the present invention, the intent recognition model includes an extraction and conversion layer, a local intent analysis layer, an intent association analysis layer, and a global intent integration layer;
[0083] In a possible embodiment, the specific steps of performing intent recognition on each speech pulse subsequence combination using the intent recognition model to obtain a second initial recognition result for each speech pulse subsequence combination include:
[0084] The extraction conversion layer extracts features from each subsequence in each speech pulse subsequence combination to obtain the converted subsequence corresponding to each subsequence;
[0085] The local intent analysis layer analyzes the feature change trends of adjacent frequency points in each converted subsequence to obtain the local intent features and feature scores of each converted subsequence;
[0086] For each sub-combination in each speech pulse sub-sequence combination, the intention association analysis layer constructs the intention association matrix of each sub-combination based on the characteristic pattern of the local intention features of each converted sub-sequence in each sub-combination;
[0087] The global intent integration layer determines the second initial intent recognition result and the second intent recognition score for each speech pulse subsequence combination based on the intent association matrix of each subcombination and the local intent features and feature scores of each converted subsequence in each subcombination.
[0088] In a possible embodiment, fusing the first initial intent recognition result with the third initial intent recognition result includes:
[0089] Deconstructing the first initial intent recognition result and the third initial intent recognition result according to preset semantic units to obtain a first semantic unit set and a second semantic unit set respectively;
[0090] Based on the semantic relevance of each element in the first semantic unit set and the second semantic unit set, the first semantic unit set and the second semantic unit set are reorganized to obtain reorganized semantic units and construct an intention relationship network;
[0091] The edges of each node in the intention relationship network are calculated based on semantic similarity and co-occurrence frequency. After multiple rounds of information propagation, fused information is obtained to obtain the target intention recognition result of the target speech pulse subsequence.
[0092] It should be noted that the intention recognition score in the third initial recognition result is greater than or equal to the intention recognition score in the first initial recognition result; if the intention screening result is an empty set, the first initial intention recognition result in the first initial recognition result is determined as the target intention recognition result for each target speech pulse subsequence.
[0093] In a possible embodiment, the process of obtaining the recombined semantic units and constructing the intention relationship network includes:
[0094] The reorganized semantic units are mapped to a preset high-dimensional semantic space to obtain the mapped high-dimensional semantic vectors. The mapped high-dimensional semantic vectors are then used as nodes to construct an intention relationship network. The node edges between each node in the intention relationship network are calculated based on semantic similarity and co-occurrence frequency to represent the strength of the association between intentions.
[0095] For each node in the intention relationship network, information propagation is performed based on the association strength of each node. After multiple rounds of propagation, the fused node information of each node is obtained; the fused node information of each node integrates the intention information of its surrounding nodes.
[0096] In one possible embodiment, the edges of each node in the intent relationship network are calculated based on semantic similarity and co-occurrence frequency, and fused information is obtained through multiple rounds of information propagation to obtain the target intent recognition result of the target speech pulse subsequence, including:
[0097] Perform information propagation and feature extraction on each node of the intention relationship network to obtain the first node intention feature of the fused node information. Based on the first node intention feature, determine the intention feature set.
[0098] Specifically, feature extraction is performed on the fused node information of each node to obtain the first node intention feature of each fused node information;
[0099] The intention feature set is determined based on the difference between the first node intention feature of each fused node information and the second node intention feature in the intention feature category corresponding to the first node intention feature, and the correlation between the first node intention feature of each fused node information and its corresponding target speech pulse subsequence.
[0100] Based on the preset semantic knowledge base, each target intent feature in the intent feature set is semantically expanded to construct a semantic feature cluster;
[0101] Perform topic mining based on semantic feature clusters to obtain intent topics;
[0102] The target intention recognition result corresponding to each target speech pulse subsequence is obtained based on the probability distribution of the intention theme.
[0103] The data push module is used to receive and splice the target intention recognition results to obtain the voice intention recognition results, and push the target power grid data corresponding to the voice intention recognition results to the user terminal.
[0104] Specifically, the speech recognition module determines the target intention recognition result corresponding to each target speech pulse subsequence based on the intention recognition result of each target speech pulse subsequence and the intention recognition result of each speech pulse subsequence combination, among which the target intention recognition results include specific intention categories such as "query power consumption", "query electricity bill", and "push electricity consumption report".
[0105] Furthermore, the data push module splices each target intent recognition result according to its order in the multimedia interactive voice to obtain the voice intent recognition result. For example, if the target intent recognition result of the first target voice pulse subsequence is "query", and the target intent recognition result of the second target voice pulse subsequence is "battery power", then the spliced voice intent recognition result is "query battery power".
[0106] Furthermore, the data push module searches the power grid data database for the corresponding target power grid data based on the voice intent recognition result. For example, if the voice intent recognition result is "query power", the user's current power data is queried to obtain the user's target power grid data.
[0107] Furthermore, the data push module encapsulates the target power grid data in a format such as JSON, and pushes it to the user's terminal device via the network.
[0108] Embodiment 2 is the second embodiment of the present invention, which differs from the first two embodiments in that:
[0109] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0110] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0111] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0112] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0113] Example 3, reference Figure 2 , which is a third embodiment of the present invention, provides a method for pushing power grid data based on multimedia interaction, including: S100: responding to a data push request initiated by a user, and collecting user images and multimedia interactive voice based on the data push request;
[0114] In an embodiment of the present invention, the power grid data push system continuously monitors the network port and waits for a data push request initiated by a user. When a user triggers a data push operation on his terminal device, the request is sent to the power grid data push system via the network.
[0115] After receiving a data push request, the power grid data push system responds to the user's data push request by sending a command to the terminal device to activate the camera and capture the user's current user image, where the user image can be understood as an image of the user's face. For example, a mobile app uses the phone's camera to capture a high-definition frontal photo and uploads it to the power grid data push system to obtain the user's user image.
[0116] At the same time, the power grid data push system instructs the terminal device to turn on its microphone and begin recording the user's voice. The user can then speak commands related to data push requests, such as "Push my electricity usage details for this month." The recording process continues until the user completes their request or reaches a preset time limit. The recorded voice is temporarily stored in the terminal device as an audio file and then uploaded to the power grid data push system, resulting in the user's multimedia interactive voice.
[0117] S102: Performing facial recognition on the user to obtain a facial recognition result; if the user passes the identity verification, performing sequence decomposition on the multimedia interactive voice to obtain multiple target voice pulse subsequences;
[0118] A voice state transition probability matrix is constructed based on the voiceprint information of the user's historical interactive speech; the matrix elements in the voice state transition probability matrix represent the state transition probability of transferring from the current voiceprint category to the next voiceprint category;
[0119] The power grid data push system preprocesses the user's historical interactive voice, removes noise and interference signals, and obtains the preprocessed interactive voice.
[0120] Furthermore, the power grid data push system converts the pre-processed interactive voice from the time domain to the frequency domain through short-time Fourier transform to obtain a spectrogram, and extracts the spectral features of each time segment on the spectrogram according to a fixed time window and frequency resolution to obtain a voiceprint feature vector.
[0121] In a possible embodiment, the short-time Fourier transform formula is:
[0122]
[0123] Where STFT(s(t),τ,f) is the preprocessed interactive speech, w(t) represents the window function, τ represents the time offset, and f represents the frequency.
[0124] Extracted voiceprint feature vector V k It can be expressed as:
[0125] V k =[STFT(s(t),τ k ,f1),STFT(s(t),τ k,f2),...,STFT(s(t),τ k ,f n )]
[0126] Among them, V k is the voiceprint feature vector, k is the kth time segment, and f represents the frequency.
[0127] Furthermore, the power grid data push system obtains the number of transitions between different voiceprint categories based on the voiceprint feature vector. For each time segment, the voice state transition probability matrix P is constructed based on the change between the voiceprint category to which it belongs and the voiceprint category to which the next time segment belongs. The matrix element p in the voice state transition probability matrix P is ij Represents the state transition probability from the current voiceprint category i to the next voiceprint category j. N ij Indicates the number of times voiceprint category i is transferred to voiceprint category j, N i Represents the total number of times voiceprint category i appears, then the matrix element p ij It can be expressed as p ij =N ij / N i , the speech state transition probability matrix can be expressed as P = [p ij ].
[0128] Perform speech analysis on the pulse sequence of multimedia interactive speech based on the speech state transition probability matrix to obtain multiple speech sequence splitting points of the multimedia interactive speech;
[0129] The pulse sequence of multimedia interactive speech is sequence-splitting according to a plurality of speech sequence splitting points to obtain a plurality of target speech pulse subsequences.
[0130] Furthermore, multiple speech sequence splitting points of the multimedia interactive speech are obtained, including:
[0131] Obtain the time interval and amplitude difference of each pulse in the pulse sequence of the multimedia interactive voice; starting from the starting pulse of the pulse sequence of the multimedia interactive voice, calculate the transition path of the voiceprint category corresponding to each pulse based on the voice state transition probability matrix and the time interval and amplitude difference of each pulse, and obtain the voice state transition path of the pulse sequence of the multimedia interactive voice;
[0132] On the speech state transition path, if the state transition probability of the voiceprint category corresponding to the target pulse position is less than a preset probability threshold, the target pulse position is determined as the speech sequence splitting point.
[0133] S104: Based on the target speech pulse subsequence, obtain the target intention recognition result; the specific process includes steps S104-1-S104-4.
[0134] S104-1: Input each target speech pulse subsequence and its corresponding speech pulse subsequence combination into the intention recognition model to obtain the first initial recognition result of each target speech pulse subsequence output by the intention recognition model, and the second initial recognition result of each speech pulse subsequence combination.
[0135] S104-2: Based on the intention recognition score in the first initial recognition result and the intention recognition score in the second initial recognition result, the second initial recognition result is screened to obtain an intention screening result.
[0136] Furthermore, the power grid data push system compares the intention recognition score in each second initial recognition result with the intention recognition score in the first initial recognition result to obtain a comparison result.
[0137] Furthermore, the power grid data push system filters the second initial recognition result based on the comparison result to obtain an intent screening result, i.e., retaining the recognition results corresponding to the intent recognition scores in the second initial recognition result that are greater than or equal to the intent recognition scores in the first initial recognition result. Therefore, the intent screening result can be an empty set or a non-empty set. Therefore, it can be understood that the intent recognition score of the third initial recognition result ultimately retained in the second initial recognition result is greater than or equal to the intent recognition score in the first initial recognition result.
[0138] S104-3: If the intention screening result includes at least one third initial recognition result, the first initial intention recognition result in the first initial recognition result and the second initial intention recognition result in the third initial recognition result are fused to obtain the target intention recognition result corresponding to each target speech pulse subsequence.
[0139] Furthermore, if it is determined that the intention screening result includes at least one third initial recognition result, the power grid data push system will fuse the first initial intention recognition result in the first initial recognition result and the second initial intention recognition result in the third initial recognition result to obtain the target intention recognition result corresponding to each target speech pulse subsequence.
[0140] S104-4: If the intention screening result is an empty set, the first initial intention recognition result in the first initial recognition result is determined as the target intention recognition result of each target speech pulse subsequence.
[0141] Furthermore, if it is determined that the intention screening result is an empty set, the power grid data push system determines the first initial intention recognition result in the first initial recognition result as the target intention recognition result of each target speech pulse subsequence.
[0142] In a possible embodiment, the specific process of step S104-1 is as follows:
[0143] After the power grid data push system inputs each target speech pulse subsequence and its corresponding speech pulse subsequence combination into the intention recognition model, the extraction conversion layer extracts features from each target speech pulse subsequence and each subsequence in the speech pulse subsequence combination, converts the target speech pulse subsequence and each subsequence in the speech pulse subsequence combination from the time domain to the frequency domain to obtain richer speech feature information, and obtains the first converted subsequence corresponding to the target speech pulse subsequence and the second converted subsequence corresponding to each subsequence in the speech pulse subsequence combination. For the speech pulse subsequence combination {V i-2 ,V i-1 ,V i}, {V i-1 ,V i ,V i+1} and {V i ,V i+1 ,V i+2}, V i-2 represents the first two segments of the target speech pulse subsequence, V i-1 represents the previous segment of the target speech pulse subsequence, V i+1 represents the next segment of the target speech pulse subsequence, V i+2 Represents the last two segments of the target speech pulse subsequence. For each segment of subsequence V n , and its corresponding transformed subsequence is:
[0144]
[0145] Among them, T(V n ) is the converted subsequence, N represents the subsequence V n The total number of sample points in; k represents the index variable, V n (k) represents the subsequence V n The kth sample value in , j represents the description unit, and a represents the preset parameter.
[0146] The local intent analysis layer analyzes the feature change trend of adjacent frequency points in each converted subsequence to obtain the local intent feature of each converted subsequence;
[0147] Furthermore, the local intent analysis layer analyzes the feature change trends of adjacent frequency points in the first converted subsequence and each second converted subsequence to obtain the local intent features of the first converted subsequence and the local intent features of each second converted subsequence. The local intent features of the first converted subsequence are the first initial intent recognition results.
[0148] In a possible embodiment, for the converted subsequence T(V n), its corresponding local intention feature L(V n ) is as follows:
[0149] L(V n )=[T(V n )(i+1)-T(V n )(i-1)] / T(V n )(i)
[0150] Among them, L(V n ) is the local intention feature, T(V n )(i+1) represents the converted subsequence T(V m ) in the i+1th frequency point feature, T(V n )(i-1) represents the converted subsequence T(V n ) in the i-1th frequency point feature, T(V n )(i) represents the converted subsequence T(V n ) is the feature of the i-th frequency point in .
[0151] Furthermore, the local intent analysis layer obtains the feature score of the first converted subsequence based on the local intent feature map of the first converted subsequence, and obtains the feature score of the second converted subsequence based on the local intent feature map of each second converted subsequence. The feature score of the first converted subsequence is the first intent recognition score.
[0152] The intent association analysis layer constructs the intent association matrix of each sub-combination based on the characteristic pattern of the local intent features of each converted sub-sequence in each sub-combination.
[0153] Furthermore, each speech pulse subsequence combination can be divided into multiple sub-combinations. For each sub-combination in each speech pulse subsequence combination, the intent association analysis layer determines the intent association between the local intent features of each converted subsequence in each sub-combination based on the characteristic pattern of the local intent features of each converted subsequence, and constructs an intent association matrix based on the intent association.
[0154] In a possible embodiment, sub-combination 1 includes sub-sequence V1, sub-sequence V2, and sub-sequence V3, and their corresponding local intent features are local intent feature L(V1), local intent feature L(V2), and local intent feature L(V3). Therefore, the intent correlation between the local intent features of any two converted sub-sequences in sub-combination 1 is:
[0155]
[0156] Among them, CV jk is the intention relevance, and L(Vk ) represents the local intention feature L(V j ) and local intention feature L(V k ), N represents the mean of the subsequence V n The total number of sample points in .
[0157] The global intent integration layer determines the second initial intent recognition result and intent recognition score of each speech pulse subsequence combination based on the intent association matrix of each subcombination and the local intent features and feature scores of each converted subsequence in each subcombination.
[0158] Furthermore, the global intent integration layer updates the feature score of each converted subsequence in each subcombination according to the intent association matrix of each subcombination to obtain the updated score of each converted subsequence in each subcombination.
[0159] Furthermore, the global intent integration layer eliminates the subsequences in each subcombination whose updated scores are less than the preset scores to obtain the final subsequence in each subcombination, and determines the average of the updated scores of each subsequence in the final subsequence as the intention recognition score of each subcombination, and splices the local intent features corresponding to the final subsequence to obtain the second initial intention recognition result of each subcombination. Therefore, it can be understood that the recognition result of each speech pulse subsequence combination includes the second initial intention recognition result and intention recognition score of each subcombination.
[0160] In a possible embodiment, the specific process of step S104-3 is as follows:
[0161] The first initial intent recognition result and the third initial intent recognition result are deconstructed according to preset semantic units, which include nouns, verbs and modifiers. The power grid data push system deconstructs the first initial intent recognition result and the second initial intent recognition result according to the semantic units of nouns, verbs and modifiers, respectively, to obtain multiple first semantic unit elements and multiple second semantic unit elements, and constructs a first semantic unit set based on the multiple first semantic unit elements, and constructs a second semantic unit set based on the multiple second semantic unit elements.
[0162] Based on the semantic relevance of each element in the first semantic unit set and the second semantic unit set, the first semantic unit set and the second semantic unit set are reorganized and the reorganized semantic units are mapped to a preset high-dimensional semantic space to obtain the mapped high-dimensional semantic vectors, and the mapped high-dimensional semantic vectors are used as nodes to construct an intention relationship network.
[0163] Furthermore, for each first semantic unit element in the first semantic unit set, the power grid data push system calculates the semantic relevance between the first semantic unit element and the second semantic unit element from two levels: a lexical level and a semantic level. The lexical level represents the morphological similarity of the semantic unit elements in terms of vocabulary, and the semantic level represents the semantic similarity of the semantic unit elements in terms of semantic features.
[0164] The specific process of obtaining word form similarity through analysis at the vocabulary level is as follows:
[0165] In a possible embodiment, the power grid data push system decomposes the first semantic unit element x into a first character sequence C x ={c x1 ,c x2 ,...,c xp}, where p represents the dimension of the first character sequence, and the second semantic unit element y is decomposed into the second character sequence C y ={c y1 ,c y2 ,...,c yq}, q represents the dimension of the second character sequence.
[0166] Perform one-hot encoding on each character c. In one possible embodiment, the character set size is N c , then the encoding vector v of each character c c is an N c -dimensional vector, where only the position corresponding to character c is 1 and the rest are 0.
[0167] Furthermore, the power grid data push system sends the first character sequence C x The corresponding encoding vector v cx Each component in the second character sequence C y The corresponding encoding vector v cy For each component in , calculate the character vector similarity sim char (v cx ,v cy ), the specific formula is as follows:
[0168]
[0169] Among them, sim char (v cx ,v cy ) is the character vector similarity, v cx (r) represents the encoding vector v cx The rth component in v cy (r) represents the encoding vector v cy The rth component in .
[0170] Furthermore, the character pair similarity matrix M is constructed based on the character vector similarity c h ar , character pair similarity matrix M c h ar The matrix elements in Initialize a matrix DP[0][0]=0, the row of the matrix DP is p+1, and the column is q+1.
[0171] For i=1 to p, DP[i][0]=DP[i-1][0]+min 1≤j≤q M ij , for j = 1 to q, DP[0][j] = DP[0][j-1] + min 1≤i≤p M ij , so for i∈[1,p], j∈[1,q], DP[i][j] is:
[0172]
[0173] The power grid data push system calculates the word form similarity sim between the first semantic unit element x and the second semantic unit element y according to the matrix DP form (x,y) is:
[0174] sim form (x,y)=DP[p][q] / max(p,q)
[0175] Among them, sim form (x,y) is the word form similarity, and DP is the matrix.
[0176] The specific process of obtaining semantic similarity through analysis of the semantic layer is as follows:
[0177] In a possible embodiment, the power grid data push system extracts the first semantic feature set F of the first semantic unit element x. x ={f x1 ,f x2 ,...,f xP}, P represents the dimension of the first semantic feature set, and the second semantic feature set F is used to extract the second semantic unit element y y ={f y1 ,f y2 ,...,f yQ}, Q represents the dimension of the second semantic feature set, and assigns a unique identifier ID(f) to each semantic feature f.
[0178] Furthermore, the power grid data push system constructs a set U containing all semantic feature identifiers. The set U can be expressed as Therefore, for the first semantic unit element x, construct its vector in the semantic feature space as W x , where for the semantic feature space vector W x , whether it contains the semantic feature value of identifier k can be expressed as:
[0179]
[0180] Among them, W x (k) is the value of the semantic feature containing the identifier k, is the first semantic feature set element, Is a unique identifier.
[0181] Similarly, construct the second semantic unit element y in the semantic feature space vector W y .
[0182] Furthermore, the power grid data push system is based on the semantic feature space vector W x and semantic feature space vector W y Whether it contains the semantic feature value with identifier k, and calculates the semantic feature space vector W x and semantic feature space vector W y The semantic feature distance d between feat (W x ,W y ), the specific calculation formula is as follows:
[0183]
[0184] Among them, d feat (W x ,W y ) is the semantic feature distance.
[0185] Furthermore, the power grid data push system is based on the semantic feature space vector W x and semantic feature space vector W y The semantic feature distance between the first semantic unit element x and the second semantic unit element y is calculated to obtain the semantic similarity sim sem (x,y) is:
[0186]
[0187] Among them, sim sem (x,y) is the semantic similarity, |F x ∩F y | represents the first semantic feature set F x and the second semantic feature set F y The number of common semantic features between |F x | represents the first semantic feature set Fx The number of elements of |F y | represents the second semantic feature set F y The number of elements, d feat (W x ,W y ) is the semantic feature distance.
[0188] Furthermore, the power grid data push system calculates the semantic relevance Sim(x,y) between the first semantic unit element and the second semantic unit element based on the word form similarity and semantic similarity. The specific formula is as follows:
[0189] Sim(x,y)=α*sim form (x,y)+(1-α)*sim sem (x,y)
[0190] Among them, Sim(x,y) is the semantic correlation, α represents the preset coefficient, sim sem (x,y) is the semantic similarity.
[0191] Furthermore, the power grid data push system recombines the first semantic unit element whose semantic relevance is greater than the preset relevance threshold θ with the second semantic unit element to obtain a recombined semantic unit L new , reorganized semantic unit L new It can be expressed as L new ={x∪y|Sim(x,y)>θ}.
[0192] Furthermore, the power grid data push system maps the reorganized semantic units to a preset high-dimensional semantic space based on a deep learning dynamic mapping model. The high-dimensional semantic space is determined based on context information, and the context information is obtained based on the data push request initiated by the user. For example, if the user queries "electricity consumption" on the terminal device, the context information is "electricity query platform".
[0193] Furthermore, the power grid data push system inputs the reorganized semantic units and context information into the deep learning dynamic mapping model, and the deep learning dynamic mapping model jointly encodes the reorganized semantic units and context information to generate a high-dimensional semantic vector.
[0194] The model structure of the deep learning dynamic mapping model is similar to an encoder with an attention mechanism, so it can adaptively focus on the parts of the semantic unit related to the context information.
[0195] In a possible embodiment, the recombined semantic unit L new It can be expressed as L new ={l1,l2,...,l M}, where M represents the reorganized semantic unit L newThe specific analysis is as follows:
[0196] Deep learning dynamic mapping model for reorganized semantic unit L new Each semantic unit l in i Embedding representation e i =Embedding(l i ), and get the embedding vector matrix E = [e 1i ,e2,...,e M ].
[0197] Furthermore, the deep learning dynamic mapping model can be used to map the context information L C Encode and get the context vector l C =EncodeContext(L C ).
[0198] Furthermore, the deep learning dynamic mapping model combines the embedding vector matrix E and the context vector l according to the attention mechanism C Calculate attention weight β i , the specific formula is:
[0199]
[0200] Among them, w, W1, W2 represent weight matrices, representing w T The bias of the weight matrix w, e j 、e i is the embedding vector, M represents the reorganized semantic unit L new dimension.
[0201] Further, calculate the weighted semantic representation m e , specifically expressed as:
[0202] Furthermore, the deep learning dynamic mapping model obtains the mapped high-dimensional semantic vector V through the fully connected layer. e , specifically: V e =W3m e +b, where W3 represents the weight matrix and b represents the bias vector.
[0203] Furthermore, the power grid data push system obtains the mapped high-dimensional semantic vector V e (i) and the high-dimensional semantic vector V after mapping e (j) The semantic similarity and co-occurrence frequency between them, and the high-dimensional semantic vector V after mapping is calculated based on the semantic similarity and co-occurrence frequency e (i) and the high-dimensional semantic vector V after mapping e (j) the association strength A ij , expressed as:
[0204] A ij =Sim(V e (i),V e (j))*log[1+fre(V e (i),V e (j)]
[0205] Among them, the semantic similarity Sim(V e (i),V e (j)) is obtained by the similarity calculation formula, the co-occurrence frequency fre(V e (i),V e (j) obtained using the same dataset.
[0206] Furthermore, the power grid data push system constructs an intention relationship network with the mapped high-dimensional semantic vectors as nodes and the correlation strength between the mapped high-dimensional semantic vectors as node edges.
[0207] The edges of each node in the intention relationship network are calculated based on semantic similarity and co-occurrence frequency. After multiple rounds of information propagation, fused information is obtained to obtain the target intention recognition result of the target speech pulse subsequence.
[0208] Furthermore, for each node in the intent-relationship network, during each round of propagation, each node transmits its own node information to adjacent nodes based on the strength of the edges connecting to it. Upon receiving this node information, the adjacent node fuses it with its own existing node information to update its own node information. After multiple rounds of propagation, each node's node information is integrated with the node information of other related nodes in the intent-relationship network to obtain the fused node information for each node, allowing each node to comprehensively reflect a more comprehensive intent relationship.
[0209] In a possible embodiment, the intention relationship network includes N e nodes, node V e (i) The node information in round t is expressed as Contains node V e (i) Various feature information related to the intention represented, A ij Represents node V e (i) and node V e (j) The correlation strength between them, therefore, the information propagation process formula is as follows:
[0210]
[0211] Among them, σ represents the activation function, such as the ReLU function; γ represents the weight coefficient, and its value range is [0,1]. Represents node V e (i) Node information in round t+1, Represents node V e (i) Node information in round t.
[0212] Assume γ = 0.3, the intention relationship network includes nodes V e (1), node V e (2), node V e (3), node V e (4), node V e (5), the association strength between nodes is:
[0213]
[0214] Initially, i.e., t=0, the node information of each node is:
[0215]
[0216] Node V e (1) as an example, the node information update for the first round, i.e., t=0→t=1, is calculated as follows:
[0217]
[0218] After the activation function σ, we get:
[0219]
[0220] Further integrate its own node information to obtain node V e The node information after fusion of (1) is:
[0221]
[0222] Therefore, after multiple rounds of propagation, node V e (1) Node information after fusion It not only contains its own initial intention features, but also integrates the node V e (2), node V e (3), node V e (5) Intent information.
[0223] Based on the fused node information of each node, the target intention recognition result corresponding to each target speech pulse subsequence is obtained.
[0224] Furthermore, the power grid data push system deeply explores the potential patterns in the fused node information, extracts features closely related to the intention from the fused information of each node, and obtains the first node intention feature that can accurately represent the node intention;
[0225] The intention feature set is determined based on the difference between the first node intention feature of each fused node information and the second node intention feature in the intention feature category corresponding to the first node intention feature, and the correlation between the first node intention feature of each fused node information and its corresponding target speech pulse subsequence.
[0226] Furthermore, the power grid data push system obtains the intention feature category corresponding to the first node intention feature of each fused node information, wherein the second node intention feature in the intention feature category can be expressed as I2={I 21 ,I 22 ,...,I 2M}.
[0227] Furthermore, the power grid data push system calculates the difference Δ(I1, I2) between the first node intention feature I1 and the second node intention feature I2. The specific formula is as follows:
[0228]
[0229] Among them, sgn represents the sign function, if I 1i -I 2i >0,sgn(I 1i -I 2i )=1, if I 1i -I 2i =0,sgn(I 1i -I 2i )=0, if I 1i -I 2i <0,sgn(I 1i -I 2i )=-1;I 1i Represents the i-th element in the first node intention feature i1, i 2i Represents the i-th element in the second node intention feature i2.
[0230] Furthermore, the power grid data push system calculates the first node intention feature I1 and its corresponding target speech pulse subsequence V t The correlation between ρ(I1,V t ), the specific formula is as follows:
[0231]
[0232] Among them, ρ(I1,V t ) represents the first node intention feature I1 and its corresponding target speech pulse subsequence V t The correlation between them, N represents the target speech pulse subsequence V t The dimension of mod represents the modulo operation, V tjRepresents the target speech pulse subsequence V t The jth element in .
[0233] Furthermore, the power grid data push system is based on the difference Δ(I1, I2) and correlation ρ(I1, V t ), determine the intention feature set C in the first node intention feature I1 I , where the intention feature set C I Satisfy C I ={I1|Δ(I1,I2)<ε1∩ρ(I1,V t )<ε2}, ε1 and ε2 represent preset thresholds.
[0234] Based on the preset semantic knowledge base, each target intent feature in the intent feature set is semantically expanded to construct a semantic feature cluster.
[0235] Furthermore, the power grid data push system performs semantic expansion on each target intent feature in the intent feature set based on the preset semantic knowledge base, that is, searches for semantic information related to the target intent feature in the semantic knowledge base, and integrates the semantic information into the target intent feature to construct a semantic feature cluster.
[0236] In a possible embodiment, the semantic knowledge base is For the intention feature set C I Each target intent feature c in Ii , in the semantic knowledge base Get the feature c of each target intention Ii The relevant semantic path set is P load ={p1,p2,...,p H}. Therefore, for each target intent feature c Ii Perform semantic expansion to obtain the intention feature set C I The expanded semantic feature set for:
[0237]
[0238] Among them, p h represents the hth semantic path, semantic path p h The semantic elements in v ht Representation and semantic elements u hr The relevant vector is determined according to the structure of the semantic knowledge base; dist(c Ii ,u hr ) represents each target intention feature c Ii With the semantic element u hr Distance in semantic space.
[0239] Furthermore, the power grid data push system calculates semantic entropy based on the expanded semantic feature set. For semantic features, the specific formula of semantic entropy is as follows:
[0240]
[0241] Among them, H(μ x ,μ y ) represents the semantic feature μ x and semantic features μ y The semantic entropy between Represented in the semantic knowledge base In the given semantic feature μ x When the semantic element Probability of occurrence; Represented in the semantic knowledge base In the given semantic feature μ y When the semantic element Probability of occurrence.
[0242] Furthermore, the power grid data push system can push the expanded semantic feature set Based on the semantic entropy H(μ x ,μ y ) to cluster and obtain semantic feature clusters
[0243] Topic mining is performed based on semantic feature clusters to obtain intention topics.
[0244] Furthermore, the power grid data push system conducts in-depth mining of semantic feature clusters to find potential topic patterns in the semantic feature clusters. That is, by analyzing the complex semantic relationships between features within the semantic feature clusters, representative intent themes are extracted. The intent themes can highly summarize the core intent information contained in the semantic feature clusters.
[0245] In one possible embodiment, the semantic feature cluster Each semantic feature cluster f t Contains multiple semantic features f tj , therefore, each semantic feature cluster f t Intentional Theme It can be expressed as:
[0246]
[0247] Among them, U t Represents the semantic feature cluster f t The number of semantic features in , P(f tj ,f tk ) represents the semantic feature f tj and semantic features ftk The co-occurrence probability between Represents the semantic feature f tj and semantic features f tk The semantic dependency strength between them is determined according to the dependency rules in the semantic knowledge base.
[0248] The target intention recognition result corresponding to each target speech pulse subsequence is obtained based on the probability distribution of the intention theme.
[0249] Furthermore, the power grid data push system matches each target speech pulse subsequence with each intention topic according to the probability distribution of the intention topic, and uses the probability information of the intention topic and the intrinsic connection between the speech pulse subsequence and the intention topic to determine the target intention recognition result that is most likely to correspond to each target speech pulse subsequence.
[0250] In one possible embodiment, the intent theme Its probability distribution is P(f t )={P(g1),P(g2),...,P(g W )}, for the target speech pulse subsequence V t , and its corresponding target intention recognition result for:
[0251]
[0252] Among them, g wl Intent theme g w The composition semantic features of L w Intent theme g w The number of semantic features in proj(V t ,g wl ) represents the target speech pulse subsequence V t In the semantic feature g wl The projection on is determined by the projection rules constructed in the semantic knowledge base.
[0253] It should be noted that the present invention can quickly and accurately obtain the target intention recognition results corresponding to each target speech pulse subsequence through the intention recognition model, and thus can quickly and accurately identify the speech intention recognition results, thereby accurately pushing the target power grid data corresponding to the speech intention recognition results to the user, avoiding the push of a large amount of useless data, improving the accuracy of power grid data push, enabling users to efficiently obtain the required power grid data, and further improving the work efficiency of power grid personnel.
[0254] S106: splicing the target intention recognition results to obtain a voice intention recognition result, and pushing the target power grid data corresponding to the voice intention recognition result to the user terminal;
[0255] Furthermore, the power grid data push system splices each target intent recognition result according to its order in the multimedia interactive voice to obtain the voice intent recognition result. For example, if the target intent recognition result of the first target voice pulse subsequence is "query", and the target intent recognition result of the second target voice pulse subsequence is "electricity", then the spliced voice intent recognition result is "query electricity".
[0256] Furthermore, the power grid data push system queries the power grid data database for the corresponding target power grid data based on the voice intent recognition result. For example, if the voice intent recognition result is "query power", the system queries the user's current power data, obtains the user's target power grid data, encapsulates the target power grid data in a format such as JSON, and pushes it to the user's terminal device via the network.
[0257] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A power grid data push system based on multimedia interaction, characterized by: It includes power grid management center, data collection module, face recognition module, voice splitting module, voice recognition module and data push module; The power grid management center is used to manage each module; The data acquisition module is used to respond to the data push request initiated by the user, collect user images and multimedia interactive voice based on the data push request, and send them to the face recognition module and the voice separation module; The face recognition module receives the user image, performs face recognition on the user, obtains the face recognition result and sends it to the voice separation module; The voice splitting module receives the face recognition result and the multimedia interactive voice, and if the user passes the identity authentication, the multimedia interactive voice is sequence-splitted to obtain multiple target voice pulse subsequences and sent to the voice recognition module; The speech recognition module receives the target speech pulse subsequence, obtains the target intention recognition result and sends it to the data push module; The data push module receives and splices the target intention recognition results to obtain a voice intention recognition result, and pushes the target power grid data corresponding to the voice intention recognition result to the user terminal.
2. The power grid data push system based on multimedia interaction according to claim 1, characterized in that: Sequencing the multimedia interactive speech to obtain multiple target speech pulse subsequences and sending them to the speech recognition module includes: Construct a speech state transition probability matrix based on the voiceprint information of the user's historical interactive speech; The matrix elements in the voice state transition probability matrix represent the state transition probability of transferring from the current voiceprint category to the next voiceprint category; Performing speech analysis on the pulse sequence of the multimedia interactive speech based on the speech state transition probability matrix to obtain a plurality of speech sequence splitting points of the multimedia interactive speech; The pulse sequence of the multimedia interactive speech is sequence-splitting according to the plurality of speech sequence splitting points to obtain a plurality of target speech pulse subsequences.
3. The multimedia interactive power grid data push system according to claim 2, characterized in that: Performing speech analysis on the pulse sequence of the multimedia interactive speech based on the speech state transition probability matrix to obtain a plurality of speech sequence splitting points of the multimedia interactive speech includes: Obtaining the time interval and amplitude difference of each pulse in the pulse sequence of the multimedia interactive voice; Starting from the starting pulse of the pulse sequence of the multimedia interactive speech, calculating the transition path of the voiceprint category corresponding to each pulse based on the voice state transition probability matrix and the time interval and amplitude difference of each pulse, to obtain the voice state transition path of the pulse sequence of the multimedia interactive speech; On the speech state transition path, if the state transition probability of the voiceprint category corresponding to the target pulse position is less than a preset probability threshold, the target pulse position is determined as the speech sequence splitting point.
4. The power grid data push system based on multimedia interaction according to claim 3, characterized in that: The speech recognition module receives a target speech pulse subsequence and obtains a target intention recognition result, including: Combining multiple continuous target speech pulse subsequences to obtain a speech pulse subsequence combination; Inputting each target speech pulse subsequence and the corresponding speech pulse subsequence combination into the intention recognition model for speech recognition, obtaining a first initial recognition result for each target speech pulse subsequence and a second initial recognition result for each speech pulse subsequence combination; The first initial recognition result includes a first initial intention recognition result and a corresponding first intention recognition score; The second initial recognition result includes a second initial intention recognition result and a corresponding second intention recognition score; Filtering the second initial recognition result based on the first intention recognition score and the second intention recognition score to obtain an intention screening result; If the intention screening result includes at least one third initial recognition result, the first initial intention recognition result and the third initial intention recognition result in the third initial recognition result are fused to obtain the target intention recognition result corresponding to each target speech pulse subsequence.
5. The power grid data push system based on multimedia interaction according to claim 4, characterized in that: Fusing the first initial intent recognition result with the third initial intent recognition result in the third initial recognition result, including: Deconstructing the first initial intent recognition result and the third initial intent recognition result according to preset semantic units to obtain a first semantic unit set and a second semantic unit set respectively; Based on the semantic relevance of each element in the first semantic unit set and the second semantic unit set, the first semantic unit set and the second semantic unit set are recombined to obtain recombined semantic units and construct an intention relationship network; The edges of each node in the intention relationship network are calculated based on semantic similarity and co-occurrence frequency. After multiple rounds of information propagation, fused information is obtained to obtain the target intention recognition result of the target speech pulse subsequence.
6. The power grid data push system based on multimedia interaction according to claim 4, characterized in that: Also includes: The intention recognition score in the third initial recognition result is greater than or equal to the intention recognition score in the first initial recognition result; If the intention screening result is an empty set, the first initial intention recognition result in the first initial recognition result is determined as the target intention recognition result of each target speech pulse subsequence.
7. The multimedia interactive power grid data push system according to claim 4, characterized in that: The nodes of the intention relationship network are calculated based on semantic similarity and co-occurrence frequency. After multiple rounds of information propagation, the fused information is obtained to obtain the target intention recognition results of the target speech pulse subsequence, including: Perform information propagation and feature extraction on each node of the intention relationship network to obtain a first node intention feature of the fused node information, and determine an intention feature set based on the first node intention feature; Performing semantic expansion on each target intent feature in the intent feature set based on a preset semantic knowledge base to construct a semantic feature cluster; Perform topic mining based on the semantic feature cluster to obtain the intended topic; The target intention recognition result corresponding to each target speech pulse subsequence is obtained based on the probability distribution of the intention theme.
8. A method for pushing power grid data based on multimedia interaction, using a power grid data pushing system based on multimedia interaction according to any one of claims 1 to 7, characterized in that: The method includes: responding to a data push request initiated by a user, and collecting user images and multimedia interactive voice based on the data push request; Perform facial recognition on the user and obtain facial recognition results; If the user passes the identity authentication, the multimedia interactive voice is sequence-splitting to obtain multiple target voice pulse subsequences; Obtain target intention recognition results based on the target speech pulse subsequence; The target intention recognition results are spliced to obtain a voice intention recognition result, and the target power grid data corresponding to the voice intention recognition result is pushed to the user terminal.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of a power grid data push system based on multimedia interaction according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a power grid data push system based on multimedia interaction according to any one of claims 1 to 7 are implemented.