Data processing method and device, storage medium and electronic equipment
Patent Information
- Application Number
- CN202211426830.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-11-15
AI Technical Summary
[0003]但是现有人机交互过程中,机器无法准确理解用户话语中的语义信息,进而无法做出正确决策,比如帮助用户评价是否要购买某直播产品,或者帮助用户评价线上面试的结果是好还是坏,等等
[0019] The data processing method, apparatus, storage medium, and electronic device provided in this application determine the first and second text vectors corresponding to the spoken text data, and determine multiple target similar text vectors based on the similar text model and the first text vector. The target similar text vectors and the first text vector have similar semantics. Then, a semantic vector is generated based on the first text vector, the multiple target similar text vectors, and the semantic analysis model. Subsequently, spoken evaluation data is generated based on the semantic vector, the second text vector, and the decision model. Thus, on the one hand, it can accurately understand the semantic information in the spoken text and provide appropriate spoken evaluation, improving the human-computer interaction experience. On the other hand, it can be applied to different task scenarios and has strong applicability.
Smart Images

Figure CN118051651B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus, storage medium and electronic device. Background Technology
[0002] Currently, with the development of big data, deep learning, and 5G networks, research directions such as image processing, natural language processing, and speech processing have emerged. The development of natural language processing has changed the way humans interact with computers, shifting the interaction from a "device-centric" approach to a "human-centric" natural language interaction approach.
[0003] However, in current human-computer interaction processes, machines cannot accurately understand the semantic information in users' speech, and therefore cannot make correct decisions, such as helping users evaluate whether to buy a live-streaming product, or helping users evaluate whether the result of an online interview is good or bad, and so on. Summary of the Invention
[0004] This application provides a data processing method, apparatus, storage medium, and electronic device that can accurately understand the semantic information in the speech text and provide appropriate speech evaluation.
[0005] This application provides a data processing method, including:
[0006] Obtain the text data of the speech to be processed;
[0007] Determine the first text vector and the second text vector based on the spoken text data;
[0008] Based on the similar text model and the first text vector, multiple target similar text vectors are determined, wherein the target similar text vectors and the first text vector have similar semantics;
[0009] A semantic vector is generated based on the first text vector, multiple target similar text vectors, and a semantic analysis model.
[0010] Based on the semantic vector, the second text vector, and the decision model, speech evaluation data corresponding to the speech text data is generated.
[0011] This application also provides a data processing apparatus, including:
[0012] The acquisition module is used to acquire the text data of the speech to be processed;
[0013] The first determining module is used to determine a first text vector and a second text vector based on the spoken text data;
[0014] The second determining module is used to determine multiple target similar text vectors based on the similar text model and the first text vector, wherein the target similar text vectors and the first text vector have similar semantics;
[0015] The first generation module is used to generate a semantic vector based on the first text vector, multiple target similar text vectors, and a semantic analysis model.
[0016] The second generation module is used to generate speech evaluation data corresponding to the speech text data based on the semantic vector, the second text vector, and the decision model.
[0017] This application also provides a computer-readable storage medium storing a plurality of instructions adapted to be loaded by a processor to execute any of the above-described data processing methods.
[0018] This application also provides an electronic device, including a coupled memory and a processor, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to perform any of the above-described data processing methods.
[0019] The data processing method, apparatus, storage medium, and electronic device provided in this application determine the first and second text vectors corresponding to the spoken text data, and determine multiple target similar text vectors based on the similar text model and the first text vector. The target similar text vectors and the first text vector have similar semantics. Then, a semantic vector is generated based on the first text vector, the multiple target similar text vectors, and the semantic analysis model. Subsequently, spoken evaluation data is generated based on the semantic vector, the second text vector, and the decision model. Thus, on the one hand, it can accurately understand the semantic information in the spoken text and provide appropriate spoken evaluation, improving the human-computer interaction experience. On the other hand, it can be applied to different task scenarios and has strong applicability. Attached Figure Description
[0020] The technical solution and other beneficial effects of this application will become apparent from the following detailed description of specific embodiments in conjunction with the accompanying drawings.
[0021] Figure 1 This is a schematic diagram illustrating an application scenario of the data processing method provided in the embodiments of this application.
[0022] Figure 2 This is a flowchart illustrating the data processing method provided in an embodiment of this application.
[0023] Figure 3 This is another schematic diagram of the data processing method provided in the embodiments of this application.
[0024] Figure 4This is a schematic diagram illustrating the workflow of the similar text model and semantic analysis model provided in the embodiments of this application.
[0025] Figure 5 This is a schematic diagram of the structure of the data processing apparatus provided in the embodiments of this application.
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0027] Figure 7 Another structural schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] This application provides a data processing method, apparatus, storage medium, and electronic device.
[0030] This data processing method can be applied to, for example... Figure 1 The hardware environment shown consists of electronic devices such as terminal 101 and server 102. Figure 1 In this embodiment, server 102 connects to terminal 101 via a network and can be used to provide services to the terminal or clients installed on the terminal (e.g., clients that can provide live-streaming e-commerce services, online interview services, or online presentation services). A database can be set up on server 102 or on other devices independent of server 102 to provide data storage services for server 102. Terminal 101 can provide a graphical user interface (GUI) to present images (e.g., live-streaming e-commerce images, interview images, or presentation images) to the user and obtain user operation commands. The aforementioned network includes, but is not limited to, wide area networks (WANs), metropolitan area networks (MANs), or local area networks (LANs), and terminal 101 is not limited to personal computers (PCs), mobile phones, tablets, etc. The data processing method provided in this application embodiment can be executed by server 102, terminal 101, or jointly by server 102 and terminal 101.
[0031] Please see Figure 2 , Figure 2This is a flowchart illustrating a data processing method provided in an embodiment of this application. This data processing method relates to the field of artificial intelligence technology and can be applied to electronic devices such as terminal devices or servers. Terminal devices may include mobile phones, tablets, laptops, etc., and may also include wearable devices, smart TVs, or other smart terminals. Servers may be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing cloud services and basic cloud computing services such as big data and artificial intelligence platforms, but are not limited to these.
[0032] Specifically, the data processing method may include the following steps S101-S105, wherein:
[0033] S101. Obtain the text data of the speech to be processed.
[0034] The spoken text data can be obtained by performing speech-text recognition on the speech data. This speech data can be speech or speech extracted from video directly collected in various task scenarios, such as speech collected by a host during a live-streaming sales event, during an online interview, or during an online speech.
[0035] In some implementations, please refer to Figure 3 Prior to step S101, the data processing method may further include steps S1011-S1013, wherein:
[0036] S1011. Obtain all voice data collected within the target time period, with each voice data corresponding to at least one user.
[0037] The start and end times of the target time period can be manually specified, such as by clicking the "Start Evaluation" and "End Evaluation" buttons. Alternatively, the system can automatically determine these times. For example, in a live-streaming e-commerce scenario, the system can automatically set the start time as when the host begins promoting a product and the end time as when the product promotion concludes. Similarly, in an online interview scenario, the system can set the start time as the user begins the interview and the end time as the end time.
[0038] S1012. Determine the target user's voice data from the voice data and use it as the target voice data.
[0039] In live-streaming e-commerce scenarios, the voice data can include the voices of multiple users, such as the live-streaming host and their assistants. In this case, the live-streaming host can be targeted as the target user. For online interview scenarios, the voice data can include the voices of multiple users, such as the interviewee and the interviewer. In this case, the interviewee can be targeted as the target user. For online presentation scenarios, the voice data can include the voices of multiple users, such as the speaker and the judges. In this case, the speaker can be targeted as the target user.
[0040] S1013. Determine the text data corresponding to the target speech data to obtain the speech text data.
[0041] One approach is to use existing speech-to-text conversion models to convert target speech data into spoken text data. These conversion models can be trained neural network models, such as Automatic Speech Recognition (ASR) models.
[0042] S102. Determine the first text vector and the second text vector based on the script text data.
[0043] In some implementations, please refer to [link / reference]. Figure 3 The above step S102 may specifically include:
[0044] S1021. Using a preset sensitive word dictionary and a preset character dictionary, extract matching sensitive words and matching characters from the speech text data, and vectorize the matching sensitive words and matching characters respectively to obtain the second text vector;
[0045] S1022. Perform sentence segmentation based on the matched sensitive words, matched characters, and text data to obtain the first text vector.
[0046] The sensitive word dictionary and character dictionary can be task-scenario based on specific requirements. Each dictionary is a database; the sensitive word dictionary is an ordered collection of sensitive words, and the character dictionary is an ordered collection of characters. Specifically, the sensitive word dictionary can store words related to politics, pornography, gambling, or drugs, as well as words with special meanings. For example, for live-streaming e-commerce, a prohibited food additive might be a sensitive word; for online job interviews, salary might be a sensitive word. The character dictionary can store characters with special meanings, such as "857" representing clubbing. Typically, different task scenarios require different dictionaries. Vectorization refers to converting Chinese characters, letters, numbers, and symbols into a format that the model can accept.
[0047] The purpose of sentence segmentation is mainly to prepare for subsequent semantic analysis model processing. It can include steps such as word segmentation, stop word removal, removal of meaningless symbols, conversion of traditional Chinese characters to simplified Chinese characters, and conversion of uppercase letters to lowercase letters. It also includes short sentence segmentation, that is, dividing a complete long sentence into multiple contextual short sentences. Usually, different short sentence segmentation models are trained for different task scenarios, and the trained short sentence segmentation models are used to perform short sentence segmentation.
[0048] S103. Based on the similar text model and the first text vector, determine multiple target similar text vectors, where the target similar text vectors and the first text vector have similar semantics.
[0049] In this context, semantic similarity refers to a high degree of similarity between two text vectors. Typically, two text vectors with similar semantics also have high semantic similarity. The similarity between two vectors can be used to determine whether they are semantically similar or dissimilar. For example, if the similarity between two vectors is greater than a similarity threshold, they can be considered to have similar semantics; otherwise, they are considered dissimilar. The similarity text model primarily adds background knowledge data to the spoken text data to facilitate more accurate semantic analysis.
[0050] In some implementations, please refer to Figure 4 , Figure 4 This diagram illustrates the workflow of the similar text model and semantic analysis model provided in this application embodiment. The similar text model includes a similar text sub-model and a scene incentive sub-model, both of which are artificial intelligence models based on neural networks. The similar text sub-model is trained using reinforcement learning based on the rewards fed back by the scene incentive sub-model. The reward is the similarity calculated by the scene incentive sub-model for two text vectors. Reinforcement learning is characterized by its ability to improve the model; that is, during the reinforcement learning process, the similar text sub-model can be continuously improved, increasing its tendency to generate correct similar text vectors, thus enabling it to generate similar text vectors more accurately. Reinforcement learning (RL), also known as reward learning, evaluation learning, or reinforcement science, is a paradigm and methodology in machine learning used to describe and solve the problem of agents learning strategies to maximize rewards or achieve specific goals during interactions with the environment. The goal of reinforcement learning is to learn the mapping from environmental states to behaviors, so that the agent's chosen behavior can obtain the maximum reward from the environment, and the external environment's evaluation of the learning system (or the overall performance of the system) is optimal in some sense.
[0051] In this context, an intelligent agent primarily refers to an entity capable of interacting with its environment and altering its state, primarily through actions. In short, an intelligent agent is a neural network model consisting of several neural network layers that map monitoring results to actions. In this application, the similar text sub-model corresponds to the intelligent agent, the scene stimulus sub-model corresponds to the environment, and the similar text vectors generated by the similar text sub-model (including the target similar text vector) can be viewed as a series of actions. Under this mechanism, using reinforcement learning to train the similar text model can ensure a high degree of similarity between the similar text vectors output by the similar text sub-model and the first text vector corresponding to the original spoken text data.
[0052] In some implementations, please refer to [link / reference]. Figure 3 and Figure 4 Specifically, step S103 may include:
[0053] S1031. Input the first text vector into the similar text sub-model to obtain multiple similar text vectors;
[0054] S1032. Use the scene activation sub-model to determine the similarity between each similar text vector and the first text vector;
[0055] S1033. For any similar text vector, if the similarity of the similar text vector is greater than a preset threshold, the similar text vector is determined to be a target similar text vector.
[0056] The similar text sub-model and the scene activation sub-model share the same deep neural network framework. The preset threshold can be determined based on the actual task scenario, such as 0.7 or 0.6. The target similar text vector is determined from the similar text vectors. For example, if there are a total of q1 similar text vectors, and q2 of them have a similarity greater than the preset threshold of 0.6, then the target similar text vector is also q2.
[0057] S104. Generate a semantic vector based on the first text vector, multiple target similar text vectors, and the semantic analysis model.
[0058] Please continue to see Figure 4 Semantic vectors are obtained through semantic analysis of the first text vector (i.e., the feature vector of the spoken text data itself) and the target similar text vector (i.e., the feature vector finally output by the similar text model). The semantic analysis model can be an artificial intelligence model based on neural networks.
[0059] In some implementations, step S104 may specifically include:
[0060] S1041. Generate a concatenated vector based on the first text vector and multiple target similar text vectors;
[0061] S1042. Input the concatenated vector into the semantic analysis model to obtain the semantic vector.
[0062] In this process, the first text vector and the target similar text vector can be concatenated to form a concatenated vector, and then the semantic analysis model can be used to perform semantic analysis on the concatenated vector.
[0063] It is easy to understand that before performing the above step S103, the scene activation sub-model, the similar text sub-model, and the semantic analysis model need to be trained in advance. Among them, the scene activation sub-model needs to be trained first, and then the similar text sub-model and the semantic analysis model are trained based on the trained scene activation sub-model. The semantic analysis model and the similar text sub-model are jointly trained.
[0064] In some implementations, the training process of the scene stimulus sub-model may include the following steps:
[0065] Multiple data pairs are identified from the background knowledge database. Each data pair includes two background knowledge data points, which are either data with similar semantics or data with different semantics.
[0066] The scenario-based stimulus sub-model is trained using multiple data pairs.
[0067] The background knowledge database stores a large amount of background knowledge data, mainly referring to the environmental data corresponding to various task scenarios. Different task scenarios typically have different background knowledge data. Two pieces of background knowledge data with similar semantics in the database can be considered as a data pair, in which case their actual similarity can be set to 1. Alternatively, two pieces of background knowledge data with different semantics can be considered as a data pair, in which case their actual similarity can be set to -1. The process of training the scenario activation sub-model using these data pairs involves using the scenario activation sub-model to predict the similarity between the two pieces of background knowledge data in each data pair, and then iteratively optimizing the network parameters of the scenario activation sub-model to continuously bring the predicted similarity closer to 1 or -1.
[0068] In some implementations, after training the scene-stimulating sub-model, the semantic analysis model and the similar text sub-model can be jointly trained; that is, the data processing method may also include:
[0069] Obtain the text sample set and the semantic label corresponding to each text sample in the text sample set;
[0070] Determine the first sample vector corresponding to each text sample;
[0071] The first sample vector is input into the similar text sub-model to obtain multiple corresponding third vectors;
[0072] Multiple target third vectors are determined from multiple third vectors using a trained scene activation sub-model;
[0073] The third vector of multiple targets and the first sample vector are fused to obtain the fourth vector;
[0074] The fourth vector is input into the semantic analysis model to obtain the predicted semantic vector;
[0075] Based on the predicted semantic vectors and semantic labels, the network parameters of the semantic analysis model and the similar text sub-model are adjusted in reverse to train the semantic analysis model and the similar text sub-model.
[0076] Similar to the actual application of the semantic analysis model and similar text sub-model in steps S1031-S1033 above, the training process also requires generating multiple similar text vectors (third vectors) corresponding to the first sample vector in the same way. The trained scene activation sub-model is then used to determine the predicted similarity between the third vector and the first sample vector. Third vectors with predicted similarity greater than a threshold are selected as target third vectors. These target third vectors are then fused with the first sample vector (e.g., through simple concatenation) to obtain the fourth vector. The semantic analysis model performs semantic analysis and prediction on the fourth vector, determining the error between the predicted semantic vector and the semantic label. Based on this error, the network parameters of the semantic analysis model and the similar text sub-model are adjusted iteratively to achieve joint training of the models.
[0077] S105. Generate speech evaluation data corresponding to the speech text data based on the semantic vector, the second text vector, and the decision model.
[0078] Different task scenarios correspond to different script evaluation data. For example, in a live-streaming e-commerce scenario, the script text data is generally an introduction to a product, and the script evaluation data could be a "true or false" rating of the product. This rating helps users make a purchase decision. In an online interview scenario, the script text data is generally the interview script, and the script evaluation data could be an "excellent," "qualified," or "failed" rating for that interview script. For failed interview scripts, it can further indicate which part of the interview process had a problem. The decision model is a deep neural network model, such as a common encoder-decoder model.
[0079] In some implementations, step S105 may specifically include:
[0080] S1051. In the second text vector, the vectorized matching sensitive words and the vectorized matching characters are weighted and summed to obtain the first vector;
[0081] S1052. Merge the first vector and the semantic vector to obtain the second vector;
[0082] S1053. The second vector is processed using a decision model to obtain the speech evaluation data corresponding to the speech text data.
[0083] The fusion can be a simple concatenation. The weights in the weighted summation can be manually set, and different weights can be set for different task scenarios. For example, if the vectorized matching sensitive word is α, the vectorized matching character is β, the corresponding weights are set to Ω1 and Ω2, and the semantic vector is S2, then the first vector S1 after weighted summation can be expressed as S1=(α*Ω1+β*Ω2), and the second vector S after concatenation can be expressed as S=S1&S2.
[0084] It is easy to understand that the decision model needs to be trained in advance. Before training the decision model, the aforementioned similar text model and semantic analysis model need to be trained in advance. Furthermore, the training sample data of the decision model can include a second sample vector and a semantic sample vector. The second sample vector can be obtained by performing sensitive word matching and character matching on the aforementioned text sample set and then quantizing it. The semantic sample vector can be obtained by processing the text sample set through the trained similar text model and semantic analysis model. The training process of the decision model is similar to the actual application process, and will not be elaborated here.
[0085] As described above, the data processing method provided in this embodiment determines the first and second text vectors corresponding to the speech text data, and determines multiple target similar text vectors based on the similar text model and the first text vector. The target similar text vectors and the first text vector have similar semantics. Then, a semantic vector is generated based on the first text vector, the multiple target similar text vectors, and the semantic analysis model. Finally, speech evaluation data is generated based on the semantic vector, the second text vector, and the decision model. This method can accurately understand the semantic information in the speech text and provide appropriate speech evaluation to improve the human-computer interaction experience. In addition, it is applicable to different task scenarios and has strong applicability.
[0086] Based on the methods described in the above embodiments, this embodiment will be further described from the perspective of a data processing device. This data processing device can be implemented as an independent entity and can be applied to electronic devices such as servers or terminal devices.
[0087] Please see Figure 5 , Figure 5 This application provides a specific description of a data processing apparatus, which may include: an acquisition module 10, a first determination module 20, a second determination module 30, a first generation module 40, and a second generation module 50, wherein:
[0088] (1) Obtain module 10
[0089] The acquisition module 10 is used to acquire the text data of the speech to be processed.
[0090] In some implementations, the acquisition module 10 is further configured to:
[0091] Acquire all voice data collected within the target time period, with each voice data corresponding to at least one user;
[0092] Identify the target user's voice data from the voice data, and use it as the target voice data;
[0093] Determine the text data corresponding to the target speech data to obtain the speech text data.
[0094] (2) First Determination Module 20
[0095] The first determining module 20 is used to determine the first text vector and the second text vector based on the speech text data.
[0096] In some implementations, the first determining module 20 is specifically used for:
[0097] Using a pre-defined sensitive word dictionary and a pre-defined character dictionary, matching sensitive words and characters are extracted from the speech text data, and the matching sensitive words and characters are vectorized to obtain a second text vector.
[0098] The first text vector is obtained by segmenting the text based on the matching sensitive words, matching characters, and textual data.
[0099] (3) Second determination module 30
[0100] The second determining module 30 is used to determine multiple target similar text vectors based on the similar text model and the first text vector, wherein the target similar text vectors and the first text vector have similar semantics.
[0101] In some implementations, the similar text model includes a similar text sub-model and a scene stimulus sub-model, and the second determining module 30 is specifically used for:
[0102] Input the first text vector into the similar text sub-model to obtain multiple similar text vectors;
[0103] The similarity between each similar text vector and the first text vector is determined using a scene-based activation sub-model.
[0104] For any similar text vector, if the similarity of the similar text vector is greater than a preset threshold, the similar text vector is determined to be a target similar text vector.
[0105] (4) First generation module 40
[0106] The first generation module 40 is used to generate a semantic vector based on the first text vector, multiple target similar text vectors, and a semantic analysis model.
[0107] In some implementations, the first generation module 40 is specifically used for:
[0108] Generate a concatenated vector based on the first text vector and multiple target similar text vectors;
[0109] The concatenated vector is input into the semantic analysis model to obtain the semantic vector.
[0110] In some embodiments, the data processing apparatus further includes a training module for:
[0111] Multiple data pairs are identified from the background knowledge database. Each data pair includes two pieces of background knowledge data, which are either data with similar semantics or have different semantics.
[0112] The scenario-based stimulus sub-model is trained using multiple data pairs.
[0113] In some implementations, the training module is also used for:
[0114] Obtain the text sample set and the semantic label corresponding to each text sample in the text sample set;
[0115] Determine the first sample vector corresponding to each text sample;
[0116] The first sample vector is input into the similar text sub-model to obtain multiple corresponding third vectors;
[0117] Multiple target third vectors are determined from multiple third vectors using a trained scene activation sub-model;
[0118] The third vector of multiple targets and the first sample vector are fused to obtain the fourth vector;
[0119] The fourth vector is input into the semantic analysis model to obtain the predicted semantic vector;
[0120] Based on the predicted semantic vectors and semantic labels, the network parameters of the semantic analysis model and the similar text sub-model are adjusted in reverse to train the semantic analysis model and the similar text sub-model.
[0121] (5) Second generation module 50
[0122] The second generation module 50 is used to generate speech evaluation data corresponding to the speech text data based on the semantic vector, the second text vector and the decision model.
[0123] In some implementations, the second generation module 50 is specifically used for:
[0124] The first vector is obtained by weighted summing of the vectorized matching sensitive words and the vectorized matching characters in the second text vector;
[0125] The first vector and the semantic vector are fused to obtain the second vector;
[0126] The second vector is processed using a decision model to obtain the speech evaluation data corresponding to the speech text data.
[0127] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.
[0128] In addition, this application also provides an electronic device, which may be a smartphone, tablet computer, or other similar device. Figure 6 As shown, the electronic device 200 includes a processor 201 and a memory 202. The processor 201 and the memory 202 are electrically connected.
[0129] The processor 201 is the control center of the electronic device 200. It connects various parts of the electronic device through various interfaces and lines. By running or loading the application program stored in the memory 202 and calling the data stored in the memory 202, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.
[0130] In this embodiment, the processor 201 in the electronic device 200 loads the instructions corresponding to the processes of one or more application programs into the memory 202 according to the following steps, and the processor 201 runs the application programs stored in the memory 202 to realize various functions:
[0131] Obtain the text data of the speech to be processed;
[0132] Determine the first text vector and the second text vector based on the script text data;
[0133] Based on the similar text model and the first text vector, multiple target similar text vectors are determined, and the target similar text vectors and the first text vector have similar semantics.
[0134] Generate semantic vectors based on the first text vector, multiple target similar text vectors, and a semantic analysis model;
[0135] Based on the semantic vector, the second text vector, and the decision model, generate the corresponding speech evaluation data for the speech text data.
[0136] Figure 7 A specific structural block diagram of an electronic device provided in an embodiment of the present invention is shown. This electronic device can be used to implement the data processing method provided in the above embodiments. The electronic device may include a smartphone or a server.
[0137] The electronic device may include a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a radio frequency (RF) circuit 303, a power supply 304, an input unit 305, and a display unit 306, etc. Those skilled in the art will understand that the electronic device structure shown in the figures does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0138] The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the electronic device. Optionally, the processor may include one or more processing cores; preferably, the processor may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor.
[0139] The memory 302 can be used to store software programs (computer programs) and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.
[0140] RF circuit 303 can be used for signal reception and transmission during information transmission and reception. Specifically, it receives downlink information from the base station and hands it over to one or more processors 301 for processing; additionally, it transmits uplink data to the base station. Typically, RF circuit 303 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 303 can also communicate wirelessly with networks and other devices. This wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System for Mobile Communications), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, and SMS (Short Messaging Service).
[0141] The electronic device also includes a power supply 304 (such as a battery) that supplies power to various components. Preferably, the power supply 304 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The power supply 304 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0142] The electronic device may also include an input unit 305, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, in one embodiment, the input unit 305 may include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display or touchpad, can collect user touch operations on or near it (e.g., user operations using fingers, styluses, or any suitable object or accessory on or near the touch-sensitive surface) and drive corresponding connection devices according to a pre-set program. Optionally, the touch-sensitive surface may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 301, and can receive and execute commands from the processor 301. Furthermore, various types of touch-sensitive surfaces, such as resistive, capacitive, infrared, and surface acoustic wave, can be used to implement the touch-sensitive surface. In addition to the touch-sensitive surface, the input unit 305 may also include other input devices. Specifically, other input devices may include, but are not limited to, one or more of the following: a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick.
[0143] The electronic device may further include a display unit 306, which can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. The display unit 306 includes multiple hardware display processing units, a video frame processing module, a display screen, etc. The multiple hardware display processing units and the video frame processing module can be integrated into a processing chip. The display screen may include a display panel, optionally configured as a liquid crystal display (LCD), an organic light-emitting diode (OLED), or similar form. Furthermore, a touch-sensitive surface may cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it transmits the information to the processor 301 to determine the type of touch event. Subsequently, the processor 301 provides corresponding visual output on the display panel according to the type of touch event. Although in the figures, the touch-sensitive surface and the display panel are implemented as two separate components to achieve input and output functions, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve input and output functions.
[0144] Although not shown, the electronic device may also include a camera, Bluetooth module, etc., which will not be described in detail here. The electronic device also includes a first splicing module, which includes a signal processing module, multiple data processing modules connected to the signal processing module, and an image splicing module connected to the multiple data processing modules. Each data processing module is connected to a corresponding first display pixel interface. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302 to achieve various functions, as follows:
[0145] Obtain the text data of the speech to be processed;
[0146] Determine the first text vector and the second text vector based on the script text data;
[0147] Based on the similar text model and the first text vector, multiple target similar text vectors are determined, and the target similar text vectors and the first text vector have similar semantics;
[0148] Generate semantic vectors based on the first text vector, multiple target similar text vectors, and a semantic analysis model;
[0149] Based on the semantic vector, the second text vector, and the decision model, generate the corresponding speech evaluation data for the speech text data.
[0150] The electronic device can implement the steps of any embodiment of the data processing method provided in the embodiments of this application. Therefore, it can achieve the beneficial effects that any data processing method provided in the embodiments of this application can achieve. For details, please refer to the previous embodiments, which will not be repeated here.
[0151] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions (computer programs), or by instructions (computer programs) controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of the present invention provide a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any embodiment of the data processing method provided by the present invention.
[0152] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0153] Since the computer program stored in the storage medium can execute the steps in any of the data processing method embodiments provided in the present invention, the beneficial effects that any of the data processing methods provided in the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0154] The data processing method, apparatus, electronic device, and storage medium provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A data processing method, characterized by, include: Obtain the text data of the speech to be processed; Determine the first text vector and the second text vector based on the spoken text data; Based on the similar text model and the first text vector, multiple target similar text vectors are determined, wherein the target similar text vectors and the first text vector have similar semantics; A semantic vector is generated based on the first text vector, multiple target similar text vectors, and a semantic analysis model. Based on the semantic vector, the second text vector, and the decision model, generate the speech evaluation data corresponding to the speech text data; The step of determining the first text vector and the second text vector based on the script text data includes: Using a preset sensitive word dictionary and a preset character dictionary, matching sensitive words and matching characters are extracted from the speech text data, and the matching sensitive words and matching characters are vectorized respectively to obtain the second text vector; The first text vector is obtained by performing sentence segmentation based on the matched sensitive words, the matched characters, and the spoken text data; The similar text model includes a similar text sub-model and a scene activation sub-model. The step of determining multiple target similar text vectors based on the similar text model and the first text vector includes: The first text vector is input into the similar text sub-model to obtain multiple similar text vectors; The similarity between each similar text vector and the first text vector is determined using the scene activation sub-model. For any of the aforementioned similar text vectors, when the similarity of the similar text vectors is greater than a preset threshold, the similar text vector is determined to be a target similar text vector.
2. The data processing method of claim 1, wherein, The step of generating speech evaluation data corresponding to the speech text data based on the semantic vector, the second text vector, and the decision model includes: The first vector is obtained by weighted summation of the vectorized matching sensitive words and the vectorized matching characters. The first vector and the semantic vector are fused to obtain the second vector; The second vector is processed using a decision model to obtain the speech evaluation data corresponding to the speech text data.
3. The data processing method of claim 1, wherein, The step of generating a semantic vector based on the first text vector, multiple target similar text vectors, and a semantic analysis model includes: A concatenated vector is generated based on the first text vector and multiple target-similar text vectors; The concatenated vector is input into the semantic analysis model to obtain the semantic vector.
4. The data processing method according to claim 3, characterized in that, Before determining multiple target similar text vectors based on the similar text model and the first text vector, the process further includes: Multiple data pairs are determined from the background knowledge database. Each data pair includes two pieces of background knowledge data, which are data with similar semantics or different semantics. The scenario excitation sub-model is trained using the multiple data pairs.
5. The data processing method according to claim 4, characterized in that, After training the scene activation sub-model, the method further includes: Obtain a text sample set and the semantic label corresponding to each text sample in the text sample set; Determine the first sample vector corresponding to each of the text samples; The first sample vector is input into the similar text sub-model to obtain multiple corresponding third vectors; Multiple target third vectors are determined from multiple third vectors using the trained scene activation sub-model; The multiple target third vectors and the first sample vector are fused to obtain a fourth vector; The fourth vector is input into the semantic analysis model to obtain the predicted semantic vector; Based on the predicted semantic vector and the semantic label, the network parameters of the semantic analysis model and the similar text sub-model are adjusted in reverse to train the semantic analysis model and the similar text sub-model.
6. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire the text data of the speech to be processed; The first determining module is used to determine a first text vector and a second text vector based on the spoken text data; The second determining module is used to determine multiple target similar text vectors based on the similar text model and the first text vector, wherein the target similar text vectors and the first text vector have similar semantics; The first generation module is used to generate a semantic vector based on the first text vector, multiple target similar text vectors, and a semantic analysis model. The second generation module is used to generate speech evaluation data corresponding to the speech text data based on the semantic vector, the second text vector, and the decision model. The first determining module is specifically used for: Using a preset sensitive word dictionary and a preset character dictionary, matching sensitive words and matching characters are extracted from the speech text data, and the matching sensitive words and matching characters are vectorized respectively to obtain the second text vector; The first text vector is obtained by performing sentence segmentation based on the matched sensitive words, the matched characters, and the spoken text data; The similar text model includes a similar text sub-model and a scene activation sub-model, and the second determining module is specifically used for: The first text vector is input into the similar text sub-model to obtain multiple similar text vectors; The similarity between each similar text vector and the first text vector is determined using the scene activation sub-model. For any of the aforementioned similar text vectors, when the similarity of the similar text vectors is greater than a preset threshold, the similar text vector is determined to be a target similar text vector.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted to be loaded by a processor to perform the data processing method of any one of claims 1 to 5.
8. An electronic device, characterized in that, The method includes a coupled memory and a processor, the memory storing a computer program, and the processor running the computer program in the memory to perform the steps of the data processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Semantic similarity model training method and device, semantic similarity recognition method and device and electronic device
CN110334177A
Representative verbal skill fragment extraction device and method based on seat voice segmentation
CN111475634A