Multimedia data auditing method and device, computer equipment, readable storage medium and program product
By text conversion and error correction processing of multimedia data in the sales and service process of financial products, and using data dictionaries to replace keywords that do not conform to the context, the problem of insufficient speech recognition capabilities is solved and the accuracy of recognition and review is improved.
Patent Information
- Application Number
- CN202411966666.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
In the dual-recorded content recognition process performed by the prior art in the process of financial product sales and services, the speech recognition ability cannot recognize words with the same or similar pronunciations, resulting in low recognition accuracy and affecting the accuracy of the review results.
By obtaining multimedia data (recording and video), identifying it into initial text data, extracting keywords and context information. If it does not match the context, use a data dictionary to replace keywords with homophones, generate target text data, and marking the key fields according to preset audit rules.
It improves the identification accuracy of multimedia data and the accuracy of audit results, and ensures the accurate identification and review of key information in the sales and service of financial products.
Smart Images

Figure CN119990106A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a multimedia data audit method, device, computer equipment, computer-readable storage medium and computer program product. Background Art
[0002] In order to safeguard the legitimate rights and interests of financial consumers and regulate the sales of financial products, audio and video recording (referred to as "double recording") is conducted during the sales and service process of financial products.
[0003] In order to meet the regulatory requirements for dual recording, financial institutions usually need to make compliance judgments on the content of dual recordings, a process called "quality inspection". Quality inspection generally uses image recognition, voice recognition and other technologies to identify the content of the collected audio and video to assist auditors in making compliance judgments.
[0004] However, since the speech recognition capability cannot recognize the context of the conversation at the time, some words with the same or similar pronunciation cannot be correctly recognized according to the application scenario, which affects the recognition accuracy and leads to inaccurate final review results. Summary of the invention
[0005] Based on this, it is necessary to provide a multimedia data processing method, apparatus, computer equipment, computer-readable storage medium and computer program product that can improve recognition accuracy in response to the above technical problems.
[0006] In a first aspect, the present application provides a multimedia data audit method, the method comprising:
[0007] Acquiring multimedia data, wherein the multimedia data includes audio data and video data;
[0008] Recognize and convert the audio and video data into initial text data; and extract a first keyword and first context information corresponding to the first keyword from the initial text data;
[0009] In the case where the first keyword does not match the first context information, the first keyword is replaced with a homophone according to a data dictionary to obtain target text data; wherein the data dictionary includes a target keyword and the homophone of the target keyword;
[0010] In the case where a key field in the target text data does not conform to a preset audit rule, the key field is marked, and the marked target text data is output, wherein the audit rule is determined based on at least one of the key fields.
[0011] In one embodiment, after the audio and video data are recognized and converted into initial text data, the method further includes:
[0012] Extracting target punctuation marks from the initial text data;
[0013] Dividing the initial text data according to the position of the target punctuation mark to obtain individual text segment data;
[0014] The extracting the first keyword and the first context information corresponding to the first keyword from the initial text data includes:
[0015] For each of the text segment data, the first keyword and the first context information corresponding to the first keyword are extracted from the text segment data.
[0016] In one embodiment, when the first keyword does not match the first context information, replacing the first keyword with the homophone according to a data dictionary to obtain target text data includes:
[0017] For each of the text segment data, when the first keyword does not match the first context information, the first keyword is replaced with the homophone according to the data dictionary, and the homophone is highlighted;
[0018] Each highlighted text segment data is combined to obtain the target text data.
[0019] In one embodiment, the audit rule includes a first audit rule and a second audit rule, wherein the first audit rule is determined based on one of the key fields; the second audit rule is determined based on at least one of a logical operation relationship between the key fields, a sequence relationship between the key fields, and a word distance relationship between the key fields;
[0020] When the key field in the target text data does not conform to the preset audit rule, marking the key field includes at least one of the following:
[0021] If the key field in the target text data matches the key field in the first audit rule, it indicates that the key field in the target text data does not comply with the preset first audit rule; the key field is marked;
[0022] If at least one of the logical operation relationship between each key field in the target text data and each key field in the second audit rule, the order relationship between each key field, and the word distance relationship between each key field hits, it indicates that the key field in the target text data does not comply with the preset second audit rule; each key field is marked.
[0023] In one embodiment, the data dictionary is established by:
[0024] Acquiring historical multimedia data; the historical multimedia data includes historical audio data and historical video data;
[0025] Identify and convert the historical audio data and historical video data into historical text data;
[0026] Extracting a second keyword and second context information corresponding to the second keyword from the historical text data according to the word frequency;
[0027] using the second keyword that does not match the second context information as a target keyword;
[0028] Obtain homophones of the target keyword, and establish a data dictionary based on the target keyword and the homophones of the target keyword.
[0029] In one embodiment, the step of identifying and converting the audio and video data into initial text data includes:
[0030] Using the acquired audio data and video data as input parameters;
[0031] The speech transcription interface is called to convert the input parameters, and the initial text data is used as the output parameter of the speech transcription interface.
[0032] In a second aspect, the present application provides a multimedia data auditing device, the device comprising:
[0033] An acquisition module, used to acquire multimedia data, wherein the multimedia data includes audio data and video data;
[0034] A conversion module, used to identify and convert the audio data and the video data into initial text data; and extract a first keyword and first context information corresponding to the first keyword from the initial text data;
[0035] An audit module, used for replacing the first keyword with a homophone according to a data dictionary to obtain target text data when the first keyword does not match the first context information; wherein the data dictionary includes a target keyword and the homophone of the target keyword;
[0036] The audit module is used to mark the key fields in the target text data when the key fields do not meet the preset audit rules, and output the marked target text data, wherein the audit rules are determined based on at least one of the key fields.
[0037] In a third aspect, the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0039] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements the steps of the above method when executed by a processor.
[0040] The above-mentioned multimedia data review method, device, computer equipment, computer-readable storage medium and computer program product first obtain multimedia data, which includes audio data and video data; secondly, convert the multimedia data into text to obtain initial text data; thirdly, perform error correction check on the initial text data to determine whether the first keyword matches the corresponding first context information, that is, whether the context information matches; if not, replace the first keyword with a homophone according to the data dictionary to obtain target text data; match the first keyword that may be erroneous with the correct homophone, thereby improving the accuracy of recognition; finally, review the target text data obtained after error correction according to preset review rules, thereby improving the accuracy of the review. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 A diagram of an application environment of a multimedia data audit method in an embodiment;
[0043] Figure 2 A schematic diagram of a process of a multimedia data audit method in one embodiment;
[0044] Figure 3 A schematic diagram of a process for dividing initial text data and extracting information in one embodiment;
[0045] Figure 4 A schematic diagram of a process for obtaining target text data in one embodiment;
[0046] Figure 5 A schematic diagram of a process for reviewing key fields in target text data in one embodiment;
[0047] Figure 6 A schematic diagram of a flow chart of a method for establishing a data dictionary in an embodiment;
[0048] Figure 7 A schematic diagram of a process of converting identification into initial text data in one embodiment;
[0049] Figure 8 A schematic diagram of a process of multimedia data review in another embodiment;
[0050] Fig. 9 A schematic diagram of a process of reviewing according to rules in one embodiment;
[0051] Fig.10 is a structural block diagram of a multimedia data auditing device in one embodiment;
[0052] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] The multimedia data audit method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The server 104 obtains multimedia data, and the multimedia data includes audio data and video data; the audio data and video data are recognized and converted into initial text data; and the first keyword and the first context information corresponding to the first keyword are extracted from the initial text data; when the first keyword does not match the first context information, the first keyword is replaced with a homophone according to the data dictionary to obtain the target text data; wherein the data dictionary includes the target keyword and the homophone of the target keyword; when the key field in the target text data does not meet the preset audit rules, the key field is marked, and the marked target text data is output, wherein the audit rules are determined based on at least one key field. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices, and the Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. The portable wearable device may be a smart watch, a smart bracelet, a head-mounted device, etc. The head-mounted device may be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0055] In an exemplary embodiment, Figure 2 As shown, a multimedia data audit method is provided, which is applied to Figure 1 The server in is taken as an example to illustrate, including the following steps S202 to S208.
[0056] in:
[0057] Step S202, acquiring multimedia data, where the multimedia data includes audio data and video data.
[0058] Among them, multimedia data includes audio data and video data.
[0059] In actual applications, in order to safeguard the legitimate rights and interests of financial consumers and regulate the sales behavior of financial products, audio and video recording (referred to as "double recording") is carried out during the sales and service process of financial products to obtain audio and video data.
[0060] Optionally, before the server acquires the multimedia data, audio data and video data of the financial product sales and service process are collected by using a recording and video recording device. The recording and video recording device communicates with the server over the network, and the recording and video recording device sends the audio data and video data to the server, which acquires the audio data and video data.
[0061] Optionally, audio and video data of the financial product sales and service process are collected by using audio and video recording equipment, and the audio and video data are uploaded to a database, and the server obtains the audio and video data from the database.
[0062] Step S204: The audio data and the video data are recognized and converted into initial text data; and a first keyword and first context information corresponding to the first keyword are extracted from the initial text data.
[0063] The context information is also the context information, the preceding text information and the following text information of the keywords, including the dialogue text between the financial practitioners and their service objects.
[0064] Optionally, the server converts the audio and video data into initial text data through automatic speech recognition (ASR), and extracts the first keyword and the first context information corresponding to the first keyword from the initial text data. For example, the server recognizes the speech of the financial practitioners and the service objects in response feedback through ASR technology, converts it into initial text data, and extracts the first keyword and the first context information corresponding to the first keyword from the initial text data.
[0065] Step S206 , when the first keyword does not match the first context information, the first keyword is replaced with a homophone according to the data dictionary to obtain target text data.
[0066] The data dictionary includes the target keyword and its homophones. The data dictionary stores the context information of the double recording scene, pre-defining the content and form that may be misrecognized due to homophones (near-phonetic) words in the speech recognition process.
[0067] Optionally, when the first keyword does not match the first context information, that is, the first keyword in the initial text data does not match the first context information corresponding to the first keyword, the first keyword may contain erroneous content that does not match the context information. The first keyword is replaced with a homophone or a near-phonetic word according to the data dictionary, and the initial text data after the replacement is the target text data.
[0068] When the first keyword in the initial text data does not match the first context information, the first keyword is highlighted. For transcription errors of the first keyword such as "product name" or "customer name", the first keyword is highlighted, and the replaced homophones or near-phonetic words can also be highlighted or marked in other ways, such as bold, wavy lines, underline, etc.
[0069] In actual applications, the context of the double-recording scenario and the homophone database are combined to detect whether the double-recording speech text contains erroneous recognition content that does not conform to the context. If erroneous content is found, it is replaced by the data dictionary or the correct content in the homophone database.
[0070] Optionally, all audit vocabulary data is extracted from the system database and stored in a dictionary type variable. Through the dictionary iterator, the input audio text is searched for possible homophones in the dictionary. Since the text content of the dual-recorded speech should conform to the context of the dual-recorded scene, if content that does not conform to the context is detected in the dual-recorded speech text, this content is replaced with content that conforms to the context.
[0071] Optionally, when the first keyword matches the first context information, the initial text data is used as the target text data for subsequent review. When the first keyword in the initial text data matches the first context information, no special processing is performed when outputting the target text data.
[0072] Step S208, when the key fields in the target text data do not conform to the preset audit rules, the key fields are marked and the marked target text data is output.
[0073] The audit rules are determined based on at least one key field. For example, the audit rules include a first audit rule and a second audit rule, wherein the first audit rule is determined based on a single key field; the single key field can be summarized and sorted based on the words that are the focus of supervision, and each key field is searched in the text target text data. If these key fields are found, it is considered that a violation point is detected.
[0074] The second audit rule is determined based on multiple key fields, such as the order in which multiple key fields appear.
[0075] Optionally, when the key fields in the target text data do not meet the preset audit rules, the server performs a quality check on the compliance of the target text data, marks the key fields, and outputs the marked target text data. The marking can be bold, wavy, underlined, highlighted, or by changing the font color.
[0076] In actual applications, when the key field does not hit the audit rules, such as inappropriate words, the screening result column is displayed in red. This part is mainly for the customer's answer words, such as requiring the customer to reply "I have known" and other information to trigger the next step of the process. If the customer's correct reply is not detected, it will be displayed in red.
[0077] In the above-mentioned multimedia data review method, first, multimedia data is obtained, and the multimedia data includes audio data and video data; secondly, the multimedia data is converted into text to obtain initial text data; thirdly, the initial text data is checked for error correction to determine whether the first keyword matches the corresponding first context information, that is, whether the context information matches; if not, the first keyword is replaced with a homophone according to the data dictionary to obtain the target text data; the first keyword that may be erroneous is matched with the correct homophone, thereby improving the accuracy of recognition; finally, the target text data obtained after error correction is reviewed according to the preset review rules, thereby improving the accuracy of the review.
[0078] In an exemplary embodiment, the initial text data is divided and information is extracted, such as Figure 3 As shown, after the audio data and video data are recognized and converted into initial text data, steps S302 to S304 are also included. Among them:
[0079] Step S302: extract target punctuation marks from the initial text data.
[0080] Optionally, the server extracts target punctuation marks from the initial text data, wherein the target punctuation marks may be a period, an exclamation mark, and the like.
[0081] Step S304, dividing the initial text data according to the positions of the target punctuation marks to obtain individual text segment data.
[0082] Optionally, taking a period as an example, the server divides the initial text data according to the position of the period to obtain various text segment data, and thereby performs an audit check on the content of each text segment data, i.e., an error correction and replacement process.
[0083] Extracting the first keyword and the first context information corresponding to the first keyword from the initial text data includes: step S306, for each text segment data, extracting the first keyword and the first context information corresponding to the first keyword from the text segment data.
[0084] For each text segment data, the server extracts the first keyword and the first context information corresponding to the first keyword from each text segment data.
[0085] In this embodiment, the initial text data is divided into text segment data according to the target punctuation marks, and the first keyword and the corresponding first context information, i.e., the context, are extracted from each text segment data. It is determined whether the first keyword is consistent with the corresponding context, and the context factor is considered, thereby improving the accuracy of recognition.
[0086] In an exemplary embodiment, Figure 4 As shown, when the first keyword does not match the first context information, the first keyword is replaced with a homophone according to the data dictionary to obtain the target text data, including steps S402 to S404.
[0087] Step S402 : for each text segment data, when the first keyword does not match the first context information, the first keyword is replaced with a homonym according to the data dictionary, and the homonym is highlighted.
[0088] Optionally, for each text segment data, when the first keyword does not match the first context information, the first keyword is replaced with a homophone or a near-homophone according to the relationship stored in the data dictionary, and the homophone or near-homophone is highlighted.
[0089] Optionally, for each text segment data, when the first keyword matches the first context information, no special mark is made in the text segment data.
[0090] Step S404, combining each highlighted text segment data to obtain target text data.
[0091] Optionally, the server combines each highlighted text segment data according to the position of the target punctuation mark to form the target text data.
[0092] In this embodiment, by verifying each text segment data, the target text data after error correction and replacement can be obtained, and the target text data is subsequently subjected to a quality inspection, thereby improving the accuracy of the quality inspection.
[0093] In an exemplary embodiment, Figure 5 As shown, the audit rules include a first audit rule and a second audit rule, wherein the first audit rule is determined based on a key field; the second audit rule is determined based on at least one of the logical operation relationship between key fields, the order relationship between key fields, and the word distance relationship between key fields; when the key field in the target text data does not meet the preset audit rules, the key field is marked, including at least one of the following steps S502 and S504. Among them:
[0094] Step S502: If the key field in the target text data matches the key field in the first audit rule, it indicates that the key field in the target text data does not comply with the preset first audit rule; the key field is marked.
[0095] The first audit rule is determined based on a key field. A single key field can be summarized and sorted based on the words that are the focus of supervision. Each key field is searched in the text target text data. If these key fields are found, it is considered that a violation point is detected.
[0096] Optionally, if the key field in the target text data matches the key field in the first audit rule, it indicates that the key field in the target text data does not comply with the preset first audit rule; the key field is marked, and the marking can be done by bolding, wavy lines, underlining, highlighting, changing the font color, etc.
[0097] Step S504, if at least one of the logical operation relationship between each key field in the target text data and each key field in the second audit rule, the order relationship between each key field, and the word distance relationship between each key field hits, it indicates that the key field in the target text data does not comply with the preset second audit rule; each key field is marked.
[0098] The second audit rule is determined based on at least one of a logical operation relationship between key fields, a sequence relationship between key fields, and a word distance relationship between key fields.
[0099] The logical operation relationships between the key fields include "and", "or" and "not".
[0100] (1) Multiple key fields appear simultaneously and continuously: define and judge using the logical operation relationship of "and". When verifying the content of the target text data, detect whether the key fields of interest appear simultaneously in the voice text.
[0101] (2) Any one of the multiple key fields appears: Use the logical operation relationship of "or" to define and judge. When verifying the content of the target text data, check whether any of the key fields of interest appear in the voice text.
[0102] (3) Multiple key fields are not allowed to appear: Use the "not" logical operation relationship to define and judge. When verifying the content of the target text data, check whether the key field of interest appears in the voice text.
[0103] Order relationship between each keyword field: Define and judge using the order relationship. For example, for the three keyword fields of credit card, installment, and handling fee, they must appear in the order of first credit card, second installment, and finally handling fee. And verify the content of the target text data according to this rule.
[0104] Word distance relationship between each keyword field: Define and judge using the distance relationship. For example, the maximum word distance of keyword fields between the character "sù" in "complaint" and the character "jīn" in "financial product" is 8 characters. When the distance between the two keyword fields of "complaint" and "financial product" in the conversation exceeds 8 characters, it will not be detected by this word relationship expression.
[0105] Optionally, if at least one of the logical operation relationship between each keyword field in the target text data and each keyword field in the second review rule, the order relationship between each keyword field, and the word distance relationship between each keyword field is hit, it indicates that the keyword field in the target text data does not conform to the preset second review rule; mark each keyword field. The marking can be done through bold, wavy line, underline, highlighting, changing font color, etc.
[0106] In this embodiment, by customizing the preset review rules, it can flexibly respond to different application scenarios and improve the applicability of the system.
[0107] In an exemplary embodiment, as Figure 6 shown, the method for establishing a data dictionary includes steps S602 to S610. Among them:
[0108] Step S602, obtain historical multimedia data; the historical multimedia data includes historical recording data and historical video data.
[0109] Optionally, the server obtains historical multimedia data; the historical multimedia data includes historical recording data and historical video data.
[0110] Step S604, identify and convert the historical recording data and historical video data into historical text data.
[0111] Optionally, the server converts the historical recording data and historical video data into historical text data through Automatic Speech Recognition (ASR).
[0112] Step S606, extract the second keyword and the corresponding second context information of the second keyword from the historical text data according to the word frequency of occurrence.
[0113] Optionally, the server analyzes and sorts professional high-frequency words in the historical multimedia data, and extracts a second keyword and second context information corresponding to the second keyword, that is, context information, from the historical text data according to the word frequency.
[0114] Step S608: taking the second keyword that does not match the second context information as a target keyword.
[0115] Optionally, the server uses the second keyword that does not match the second context information as the target keyword.
[0116] Step S610, obtaining homophones of the target keyword, and establishing a data dictionary based on the target keyword and the homophones of the target keyword.
[0117] Optionally, the server obtains homophones or near-phonetic words of the target keyword, and establishes a data dictionary according to the target keyword and the homophones or near-phonetic words of the target keyword.
[0118] In this embodiment, the acquired text information is compared with the predefined content and form that are incorrectly recognized due to homophones (or near-homophones). If any content expression that does not conform to the context appears, it is replaced according to the content expression in the correct context, thereby improving the accuracy of recognition.
[0119] In an exemplary embodiment, Figure 7 As shown, the audio data and video data are recognized and converted into initial text data, including steps S702 to S704. Among them:
[0120] Step S702: taking the acquired audio data and video data as input parameters.
[0121] Optionally, the audio and video recording device collects the audio of the conversation between the financial practitioner and the server object, and uses the audio data as an input parameter.
[0122] Step S604, calling the speech transcription interface, converting the input parameters, and using the initial text data as the output parameters of the speech transcription interface.
[0123] Optionally, the server converts the input parameters through the speech transcription interface API and returns the output parameters, which are the initial text data completed by the speech-to-text conversion.
[0124] In this embodiment, the speech can be converted into text through the speech transcription interface API.
[0125] In an exemplary embodiment, Figure 8As shown, audio data and video data of the financial product sales and service process are collected by using a recording and video recording device. The recording and video recording device communicates with the server through a network, and the recording and video recording device sends the audio data and video data to the server, which obtains the audio data and video data.
[0126] The server converts the audio and video data into initial text data through Automatic Speech Recognition (ASR).
[0127] The server extracts the target punctuation from the initial text data. Among them, the target punctuation can be a period, an exclamation mark, etc. Taking the period as an example, the server divides the initial text data according to the position of the period to obtain each text segment data. And the content of each text segment data is audited and verified, that is, the process of error correction and replacement. For each text segment data, the server extracts the first keyword and the first context information corresponding to the first keyword from each text segment data. To perform content audit of the initial text data: the first keyword and the first context information corresponding to the first keyword, that is, the context information is consistent. If the first keyword does not match the first context information, that is, the first keyword in the initial text data does not match the first context information corresponding to the first keyword, the first keyword may contain erroneous content that does not match the context information. According to the data dictionary, the first keyword is replaced with a homophone or a near-phonetic word, and the initial text data after replacement is also the target text data. When the first keyword in the initial text data does not match the first context information, the first keyword is highlighted. For transcription errors of the first keyword, such as "product name" or "customer name", they are highlighted, and the replaced homophones or near-homophones can also be highlighted or marked in other ways. Other marks may include bold, wavy lines, underline, etc. When the first keyword matches the first context information, the initial text data is used as the target text data for subsequent review. When the first keyword in the initial text data matches the first context information, no special processing is performed when outputting the target text data. The server combines each highlighted text segment data according to the position of the target punctuation mark to form the target text data.
[0128] Among them, the data dictionary is established in such a way that the server obtains historical multimedia data; the historical multimedia data includes historical audio data and historical video data. The historical audio data and historical video data are converted into historical text data through Automatic Speech Recognition (ASR). The professional high-frequency vocabulary in the historical multimedia data is analyzed and sorted. The second keyword and the second context information corresponding to the second keyword, that is, the context information, are extracted from the historical text data according to the word frequency. The second keyword that does not conform to the second context information is used as the target keyword. The homophone or near-phonetic word of the target keyword is obtained, and the data dictionary is established according to the target keyword and the homophone or near-phonetic word of the target keyword.
[0129] The server conducts quality inspection on the target text data according to the audit rules through the dual-recording voice quality inspection engine, that is, audits the target text data; if the key fields in the target text data do not meet the preset audit rules, the server conducts quality inspection on the compliance of the target text data, marks the key fields, and outputs the marked target text data. The marking can be bold, wavy, underlined, highlighted, changing the font color, etc. If the key fields in the target text data meet the preset audit rules, the processing is completed and the output quality inspection is passed.
[0130] Among them, the process of quality inspection or review of the target text data by the dual-recording voice quality inspection engine according to the review rules, such as Fig. 9 The second audit rule includes the logical operation relationship between the key fields, the order relationship between the key fields, and the word distance relationship between the key fields.
[0131] The logical operation relationships between the key fields include "and", "or" and "not".
[0132] (1) Multiple key fields appear simultaneously and continuously: define and judge using the logical operation relationship of "and". When verifying the content of the target text data, detect whether the key fields of interest appear simultaneously in the voice text.
[0133] (2) Any one of the multiple key fields appears: Use the logical operation relationship of "or" to define and judge. When verifying the content of the target text data, check whether any of the key fields of interest appear in the voice text.
[0134] (3) Multiple key fields are not allowed to appear: Use the "not" logical operation relationship to define and judge. When verifying the content of the target text data, check whether the key field of interest appears in the voice text.
[0135] Order relationship between each keyword field: It is defined and judged using the order relationship. For example, for the three keyword fields of credit card, installment, and handling fee, they must appear in the order of first credit card, second installment, and last handling fee. And the content of the target text data is verified according to this rule.
[0136] Word distance relationship between each keyword field: It is defined and judged using the distance relationship. For example, the maximum word distance of keyword fields between the character "sù" in "complaint" and the character "jīn" in "financial product" is 8 characters. When the distance between the two keyword fields of "complaint" and "financial product" in the conversation exceeds 8 characters, it will not be detected by this word relationship expression.
[0137] As Fig. 9 As shown, the server obtains the second review rule. First, it judges whether the logical operation relationship between each keyword field is hit; second, it judges the order relationship between each keyword field; finally, it judges the word distance relationship between each keyword field.
[0138] If at least one of the order relationship between each keyword field in the target text data and the second review rule and the word distance relationship between each keyword field is hit, it indicates that the keyword field in the target text data does not conform to the preset second review rule; mark each keyword field. The marking can be done through bold, wavy line, underline, highlighting, changing font color, etc.
[0139] If none of the order relationship between each keyword field in the target text data and the second review rule and the word distance relationship between each keyword field is hit, it is regarded as passing the quality inspection.
[0140] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps do not necessarily execute in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily execute at the same moment, but can execute at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0141] Based on the same inventive concept, the embodiment of the present application also provides a multimedia data audit device for implementing the multimedia data audit method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more multimedia data audit device embodiments provided below can refer to the limitations of the multimedia data audit method above, and will not be repeated here.
[0142] In an exemplary embodiment, Fig.10 As shown, a multimedia data auditing device is provided, including: an acquisition module 1001, a conversion module 1002, an audit module 1003 and an audit module 1004, wherein:
[0143] The acquisition module 1001 is used to acquire multimedia data, where the multimedia data includes audio data and video data.
[0144] The conversion module 1002 is used to identify and convert the audio data and the video data into initial text data; and extract the first keyword and the first context information corresponding to the first keyword from the initial text data.
[0145] The audit module 1003 is used to replace the first keyword with a homophone according to a data dictionary to obtain target text data when the first keyword does not match the first context information; wherein the data dictionary includes the target keyword and homophones of the target keyword.
[0146] The audit module 1004 is used to mark the key fields in the target text data when the key fields do not meet the preset audit rules, and output the marked target text data, wherein the audit rules are determined based on at least one key field.
[0147] In an exemplary embodiment, the multimedia data auditing device further includes a segmentation module for extracting target punctuation marks from initial text data; and segmenting the initial text data according to the positions of the target punctuation marks to obtain individual text segment data.
[0148] The conversion module 1002 is further configured to extract, for each text segment data, a first keyword and first context information corresponding to the first keyword from the text segment data.
[0149] In an exemplary embodiment, the audit module 1003 is also used to replace the first keyword with a homonym according to the data dictionary and highlight the homonym for each text fragment data when the first keyword does not match the first context information; and combine each highlighted text fragment data to obtain the target text data.
[0150] In an exemplary embodiment, the audit rules include a first audit rule and a second audit rule, wherein the first audit rule is determined based on a key field; the second audit rule is determined based on at least one of a logical operation relationship between key fields, a sequence relationship between key fields, and a word distance relationship between key fields; the audit module 1004 is also used to mark the key field if the key field in the target text data matches the key field in the first audit rule, indicating that the key field in the target text data does not comply with the preset first audit rule.
[0151] The audit module 1004 is also used to mark each key field if at least one of the logical operation relationship between each key field in the target text data and each key field in the second audit rule, the order relationship between each key field, and the word distance relationship between each key field is hit, indicating that the key field in the target text data does not comply with the preset second audit rule;
[0152] In an exemplary embodiment, the multimedia data auditing device also includes a data dictionary building module, which is used to obtain historical multimedia data; the historical multimedia data includes historical audio data and historical video data; the historical audio data and historical video data are identified and converted into historical text data; a second keyword and second context information corresponding to the second keyword are extracted from the historical text data according to the frequency of occurrence of the word; the second keyword that does not conform to the second context information is used as a target keyword; homophones of the target keyword are obtained, and a data dictionary is established based on the target keyword and the homophones of the target keyword.
[0153] In an exemplary embodiment, the conversion module 1002 is further used to use the acquired audio data and video data as input parameters; call the speech transcription interface, convert the input parameters, and use the initial text data as the output parameters of the speech transcription interface.
[0154] Each module in the multimedia data audit device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0155] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig.11As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store multimedia data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multimedia data audit method is implemented.
[0156] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0157] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0158] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0159] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0160] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0161] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0162] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A multimedia data audit method, characterized in that: The method comprises: Acquiring multimedia data, wherein the multimedia data includes audio data and video data; Recognize and convert the audio and video data into initial text data; and extract a first keyword and first context information corresponding to the first keyword from the initial text data; In the case where the first keyword does not match the first context information, the first keyword is replaced with a homophone according to a data dictionary to obtain target text data; wherein the data dictionary includes a target keyword and the homophone of the target keyword; In the case where a key field in the target text data does not conform to a preset audit rule, the key field is marked, and the marked target text data is output, wherein the audit rule is determined based on at least one of the key fields.
2. The method according to claim 1, characterized in that After the audio and video data are recognized and converted into initial text data, the method further includes: Extracting target punctuation marks from the initial text data; Dividing the initial text data according to the position of the target punctuation mark to obtain individual text segment data; The extracting the first keyword and the first context information corresponding to the first keyword from the initial text data includes: For each of the text segment data, the first keyword and the first context information corresponding to the first keyword are extracted from the text segment data.
3. The method according to claim 2, characterized in that The step of replacing the first keyword with the homophone according to a data dictionary to obtain target text data when the first keyword does not match the first context information includes: For each of the text segment data, when the first keyword does not match the first context information, the first keyword is replaced with the homophone according to the data dictionary, and the homophone is highlighted; Each highlighted text segment data is combined to obtain the target text data.
4. The method according to claim 1, characterized in that The audit rule includes a first audit rule and a second audit rule, wherein the first audit rule is determined based on one of the key fields; the second audit rule is determined based on at least one of a logical operation relationship between the key fields, a sequence relationship between the key fields, and a word distance relationship between the key fields; When the key field in the target text data does not conform to the preset audit rule, marking the key field includes at least one of the following: If the key field in the target text data matches the key field in the first audit rule, it indicates that the key field in the target text data does not comply with the preset first audit rule; the key field is marked; If at least one of the logical operation relationship between each key field in the target text data and each key field in the second audit rule, the order relationship between each key field, and the word distance relationship between each key field hits, it indicates that the key field in the target text data does not comply with the preset second audit rule; each key field is marked.
5. The method according to claim 1, characterized in that The data dictionary is established in a manner including: Acquiring historical multimedia data; the historical multimedia data includes historical audio data and historical video data; Identify and convert the historical audio data and historical video data into historical text data; Extracting a second keyword and second context information corresponding to the second keyword from the historical text data according to the word frequency; using the second keyword that does not match the second context information as a target keyword; Obtain homophones of the target keyword, and establish a data dictionary based on the target keyword and the homophones of the target keyword.
6. The method according to claim 1, characterized in that The step of identifying and converting the audio and video data into initial text data includes: Using the acquired audio data and video data as input parameters; The speech transcription interface is called to convert the input parameters, and the initial text data is used as the output parameter of the speech transcription interface.
7. A multimedia data auditing device, characterized in that: The device comprises: An acquisition module, used to acquire multimedia data, wherein the multimedia data includes audio data and video data; A conversion module, used to identify and convert the audio data and the video data into initial text data; and extract a first keyword and first context information corresponding to the first keyword from the initial text data; An audit module, used for replacing the first keyword with a homophone according to a data dictionary to obtain target text data when the first keyword does not match the first context information; wherein the data dictionary includes a target keyword and the homophone of the target keyword; The audit module is used to mark the key fields in the target text data when the key fields do not meet the preset audit rules, and output the marked target text data, wherein the audit rules are determined based on at least one of the key fields.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.