Live data streaming method, device, computer equipment and storage medium

By obtaining and processing live data streams during the live broadcast process, extracting and filtering sensitive text data, and performing audio data silencing processing, the problem of poor protection of sensitive information during the live broadcast process is solved, real-time and effective protection of sensitive information is achieved.

CN115002508BActive Publication Date: 2025-06-06INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210634527.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2025-06-06
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

During the live broadcast process, traditional methods cannot effectively protect sensitive information, resulting in the leakage of sensitive information such as customer personal information and financial policies.

Method used

By obtaining the live data stream, using the text conversion network to convert the audio data into a text sequence, extract sensitive text data based on the preset sensitive text extraction strategy, and filter through the sensitive information filter network, and finally silencing the audio data stream, generating a new live data stream to send to the client.

Benefits of technology

Real-time processing and protection of sensitive information during live broadcast is realized, and the protection effect of sensitive information during live broadcast is improved, and the leakage of sensitive information is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115002508B_ABST
    Figure CN115002508B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, computer equipment and storage medium for processing live data streams. The present application relates to the field of artificial intelligence. The method comprises: obtaining a live data stream during a live broadcast; determining the text sequence corresponding to the audio data stream and the audio data corresponding to each text data in the text sequence in the audio data stream through a text conversion network; extracting multiple groups of initial sensitive text data from each text data based on a preset sensitive text extraction strategy, and inputting the multiple groups of initial sensitive text data into a sensitive information screening network to obtain each sensitive text data group; performing silencing processing on the target audio data corresponding to the sensitive text data in each sensitive text data group in the audio data stream to obtain a new audio data stream; determining a new live data stream based on the new audio data stream and the image data stream, and sending the new live data stream to the client. Thereby improving the protection effect of sensitive information during live broadcast.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a live broadcast privacy protection method, apparatus, computer equipment, storage medium and computer program product. Background Art

[0002] With the development of Internet live broadcast technology, in order to popularize and promote financial security knowledge among the whole people, banks will popularize education to the public through live broadcast. However, due to the particularity of the industry, some financial data and financial policies may be involved in the live broadcast process. Failure to pay attention may lead to the leakage of sensitive information such as customer personal information and financial policies to be released.

[0003] The traditional method of preventing the leakage of sensitive information is to edit, mute and other technical processing of the recorded content during the live broadcast before publishing it to avoid further spread of sensitive information. However, this method cannot process sensitive information during the live broadcast, resulting in poor protection of sensitive information during the live broadcast. Summary of the invention

[0004] Based on this, it is necessary to provide a live data stream method, apparatus, computer equipment, computer-readable storage medium and computer program product to address the above technical issues.

[0005] In a first aspect, the present application provides a method for processing a live data stream. The method comprises:

[0006] During the live broadcast process, a live broadcast data stream is obtained; the live broadcast data stream includes an image data stream and an audio data stream;

[0007] Determine, through a text conversion network, a text sequence corresponding to the audio data stream, and audio data corresponding to each text data in the text sequence in the audio data stream;

[0008] Based on a preset sensitive text extraction strategy, extract multiple groups of initial sensitive text data from each of the text data, and input the multiple groups of initial sensitive text data into a sensitive information screening network to obtain each sensitive text data group;

[0009] Silencing the target audio data corresponding to the sensitive text data in each of the sensitive text data groups in the audio data stream to obtain a new audio data stream;

[0010] A new live data stream is determined according to the new audio data stream and the image data stream, and the new live data stream is sent to the client.

[0011] Optionally, the extracting multiple groups of initial sensitive text data from each of the text data based on a preset sensitive text extraction strategy includes:

[0012] In each of the text data, extract each digital data in the text data by using a digital feature extraction algorithm;

[0013] The continuous digital data are stored in the same initial sensitive text data group to obtain multiple groups of the initial sensitive text data.

[0014] Optionally, the extracting multiple groups of initial sensitive text data from each of the text data based on a preset sensitive text extraction strategy includes:

[0015] By using a partitioning layer of a text extraction network, each non-numeric data in the text sequence is divided into a plurality of groups of non-numeric data;

[0016] Determining the category of each group of the non-numeric data through the recognition layer of the text extraction network;

[0017] The non-numeric data corresponding to the preset category is used as the initial sensitive text data to obtain multiple groups of initial sensitive text data.

[0018] Optionally, the non-numeric data in the text sequence is divided into a plurality of non-numeric data groups by a division layer of a text extraction network, including:

[0019] For each non-numeric data, judging whether the correlation between the non-numeric data and each non-numeric data adjacent to the non-numeric data reaches a preset correlation through the partitioning layer of the text extraction network;

[0020] When the correlation between the non-digital data and each non-digital data adjacent to the non-digital data is greater than a preset correlation, the non-digital data and each non-digital data adjacent to the non-digital data are stored in the same non-digital data group to obtain multiple non-digital data groups.

[0021] Optionally, the inputting the multiple groups of initial sensitive text data into a sensitive information screening network to obtain each sensitive text data group includes:

[0022] For each initial sensitive text data group, a confidence evaluation operation is performed on each initial sensitive text data in the initial sensitive text data group in the text sequence through a confidence evaluation network to obtain the confidence of the initial sensitive text data group;

[0023] In the case that the confidence is greater than a preset confidence threshold, the initial sensitive text data group greater than the preset confidence threshold is used as a sensitive text data group to obtain each of the sensitive text data groups.

[0024] Optionally, muting the target audio data corresponding to the sensitive text data in each of the sensitive text data groups in the audio data stream to obtain a new audio data stream includes:

[0025] In each of the audio data, select the target audio data corresponding to each of the sensitive text data in the sensitive text data group;

[0026] The target audio data is replaced with preset audio data to obtain a new audio data stream.

[0027] In a second aspect, the present application also provides a live data stream processing device. The device comprises:

[0028] An acquisition module is used to acquire a live broadcast data stream during the live broadcast process; the live broadcast data stream includes an image data stream and an audio data stream;

[0029] An extraction module, used to determine, through a text conversion network, a text sequence corresponding to the audio data stream, and audio data corresponding to each text data in the text sequence in the audio data stream;

[0030] A screening module, configured to extract multiple groups of initial sensitive text data from each of the text data based on a preset sensitive text extraction strategy, and input the multiple groups of initial sensitive text data into a sensitive information screening network to obtain each sensitive text data group;

[0031] a silencing module, configured to perform silencing processing on target audio data corresponding to the sensitive text data in each of the sensitive text data groups in the audio data stream to obtain a new audio data stream;

[0032] The sending module is used to determine a new live data stream according to the new audio data stream and the image data stream, and send the new live data stream to the client.

[0033] Optionally, the screening module is specifically used for:

[0034] In each of the text data, extract each digital data in the text data by using a digital feature extraction algorithm;

[0035] The continuous digital data are stored in the same initial sensitive text data group to obtain multiple groups of the initial sensitive text data.

[0036] Optionally, the screening module is specifically used for:

[0037] By using a partitioning layer of a text extraction network, each non-numeric data in the text sequence is divided into a plurality of groups of non-numeric data;

[0038] Determining the category of each group of the non-numeric data through the recognition layer of the text extraction network;

[0039] The non-numeric data corresponding to the preset category is used as the initial sensitive text data to obtain multiple groups of initial sensitive text data.

[0040] Optionally, the screening module is specifically used for:

[0041] For each non-numeric data, judging whether the correlation between the non-numeric data and each non-numeric data adjacent to the non-numeric data reaches a preset correlation through the partitioning layer of the text extraction network;

[0042] When the correlation between the non-digital data and each non-digital data adjacent to the non-digital data is greater than a preset correlation, the non-digital data and each non-digital data adjacent to the non-digital data are stored in the same non-digital data group to obtain multiple non-digital data groups.

[0043] Optionally, the screening module is specifically used for:

[0044] For each initial sensitive text data group, a confidence evaluation operation is performed on each initial sensitive text data in the initial sensitive text data group in the text sequence through a confidence evaluation network to obtain the confidence of the initial sensitive text data group;

[0045] In the case that the confidence is greater than a preset confidence threshold, the initial sensitive text data group greater than the preset confidence threshold is used as a sensitive text data group to obtain each of the sensitive text data groups.

[0046] Optionally, the muffler module is specifically used for:

[0047] In each of the audio data, select the target audio data corresponding to each of the sensitive text data in the sensitive text data group;

[0048] The target audio data is replaced with preset audio data to obtain a new audio data stream.

[0049] In a third aspect, the present application provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the methods in the first aspect are implemented.

[0050] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of any one of the methods in the first aspect are implemented.

[0051] In a fifth aspect, the present application provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of any one of the methods in the first aspect are implemented.

[0052] The above-mentioned sensitive information protection method, device, computer equipment, storage medium and computer program product during live broadcast, obtains a live broadcast data stream during the live broadcast process; the live broadcast data stream includes an image data stream and an audio data stream; through a text conversion network, determines the text sequence corresponding to the audio data stream, and the audio data corresponding to each text data in the text sequence in the audio data stream; based on a preset sensitive text extraction strategy, extracts multiple groups of initial sensitive text data from each of the text data, and inputs the multiple groups of initial sensitive text data into a sensitive information screening network to obtain each sensitive text data group; mutes the target audio data corresponding to the sensitive text data in each of the sensitive text data groups in the audio data stream to obtain a new audio data stream; determines a new live broadcast data stream based on the new audio data stream and the image data stream, and sends the new live broadcast data stream to the client. By muting the audio data involving sensitive information in the audio data stream during live broadcast to obtain a new audio data stream, so that the new audio data stream does not contain sensitive information, and sends the new audio data stream to the client, the protection effect of sensitive information during live broadcast is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A schematic diagram of a flow chart of a live data stream processing method in one embodiment;

[0054] Figure 2 A schematic diagram of a flow chart of a method for screening each sensitive text data group in one embodiment;

[0055] Figure 3 A schematic diagram of a flow chart of a live data stream example in an embodiment;

[0056] Figure 4 is a structural block diagram of a live data stream processing device in an embodiment;

[0057] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] The live data stream processing method provided by the embodiment of the present application can be applied to a terminal, a server, or a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal may include but is not limited to various personal computers, laptops, tablet computers, etc. The terminal first converts the audio data stream in the video information obtained during the live broadcast in real time into a text sequence, and extracts the initial sensitive text data group from the text sequence. Secondly, through the sensitive information screening network, in the initial sensitive text data group obtained by the first screening, each sensitive text data group is screened again. Finally, after the audio data corresponding to the text data in each sensitive text data group in the audio data stream is muted, it is recombined with the acquired image data stream to obtain a new live data stream, and the new live data stream is sent to the client. Thereby improving the protection effect of sensitive information during live broadcast.

[0060] In one embodiment, Figure 1 As shown, a live data stream processing method is provided, which is described by taking the method applied to a terminal as an example, and includes the following steps:

[0061] Step S101, during the live broadcast process, obtain the live broadcast data stream.

[0062] The live data stream includes an image data stream and an audio data stream.

[0063] In this embodiment, during the live broadcast process, the terminal periodically obtains the live broadcast data stream within a preset time from the live broadcast end, and divides the current live broadcast data stream into an audio data stream and an image data stream. The fixed time period is the delay period when the live broadcast data stream is transmitted from the live broadcast end to the client during the live broadcast process. After the terminal obtains the live broadcast data stream, it stores the live broadcast data stream in a preset buffer area of ​​the terminal.

[0064] The preset duration may be, but is not limited to, 1 second, 2 seconds, 5 seconds, etc.

[0065] Step S102, determining the text sequence corresponding to the audio data stream and the audio data corresponding to each text data in the text sequence in the audio data stream through a text conversion network.

[0066] In this embodiment, the terminal extracts the acquired audio data stream from the buffer area, and recognizes the audio data stream through a speech recognition network (i.e., a text conversion network). Then, the terminal converts the audio data stream into a text sequence, wherein the text sequence contains a plurality of text data, and the arrangement order of the text data is the same as the order of the audio data stream. After obtaining the text sequence, the terminal sets a timestamp according to the start and end time points of the time period of the audio data stream corresponding to each text data, and obtains a section of audio data, wherein the text data corresponds to the audio data one by one. The speech recognition network can be, but is not limited to, speech recognition technology (Automatic Speech Recognition, ASR).

[0067] When identical text data appears in the text data, the terminal marks the identical text data from front to back in the order of the text sequence, and at the same time, when marking the audio data corresponding to the text data, the corresponding mark is also added to the audio data. The marking method can be but is not limited to marking by adding a timestamp in the audio data stream.

[0068] For example, if the text sequence is "Today's weather is very good.", the terminal marks the first "day" in the text sequence as the first "day", such as "day 1 ", and mark the audio data corresponding to the first "day" as the audio data of the first "day", that is, add the timestamp of the first "day" to the start time point and end time point of the audio data corresponding to the first "day", and the terminal marks the "day" after the text sequence as the second "day", such as "day 2 ", and mark the audio data corresponding to the second "day" as the audio data of the second "day", that is, add the first "day" timestamp to the start time point and end time point of the audio data corresponding to the first "day".

[0069] For another example, the total length of the audio data stream is 2 seconds. After the terminal recognizes the audio data stream through the speech recognition network, the text sequence obtained is "Today's weather is very good." Then the text sequence contains 8 text data. Among them, the starting time point of the audio data stream corresponding to the text data "today" is 0.0 seconds, and the end time point is 0.2 seconds. The terminal marks the audio data stream from 0.0 seconds to 0.2 seconds as the corresponding audio data of the text data "today"; the starting time point of the audio data stream corresponding to the second text data "day" in the text sequence is 1.0 seconds, and the end time point is 1.2 seconds. The terminal marks the audio data stream from 1.0 seconds to 1.2 seconds as the audio data corresponding to the second "day". Similarly, the terminal divides the audio data corresponding to each text data, and then the audio data stream can be divided into 8 audio data.

[0070] Step S103, based on a preset sensitive text extraction strategy, extract multiple groups of initial sensitive text data from each text data, and input the multiple groups of initial sensitive text data into a sensitive information screening network to obtain various sensitive text data groups.

[0071] In this embodiment, the terminal presets a sensitive text extraction strategy, and performs an extraction operation on the obtained text sequence according to the sensitive text extraction strategy, extracting multiple groups of initial sensitive text data from each text data in the text sequence. The specific extraction operation will be described in detail later.

[0072] The terminal screens each group of initial sensitive text data through the sensitive information screening network to obtain multiple groups of sensitive text data, each group of sensitive text data contains multiple sensitive text data. The specific screening operation will be described in detail later.

[0073] The sensitive information screening network can be, but is not limited to, Bidirectional Encoder Representations from Transformers (BERT).

[0074] Step S104, muting the target audio data corresponding to the sensitive text data in each sensitive text data group in the audio data stream to obtain a new audio data stream.

[0075] In this embodiment, the terminal searches for audio data corresponding to each sensitive text data in the sensitive text data group according to the tracing back operation marked in step S102, and uses the audio data corresponding to the sensitive text data as the target audio data. The terminal mutes each target audio data, and replaces the original target audio data in the audio data stream with the muted target audio data to obtain a new audio data stream. The specific muting process will be described in detail later.

[0076] Step S105, determining a new live data stream according to the new audio data stream and the image data stream, and sending the new live data stream to the client.

[0077] In this embodiment, the terminal merges the new audio data stream and the image data stream corresponding to the new audio data stream, and uses the merged data stream as the new live data stream. The terminal sends the new live data stream directly to the client, deletes the original live data stream from the cache area, and prepares to obtain the live data stream for a fixed time period of the current live broadcast.

[0078] Specifically, the terminal splices the starting time point of the new audio data stream with the starting time point of the image data stream, and splices the ending time point of the new audio data stream with the ending time point of the image data stream to obtain a new live data stream. The audio data involving sensitive information in the new live data stream will be modified, thereby avoiding the leakage of sensitive information. The specific splicing method can be but is not limited to any audio and image merging and cutting method in the video cutting technology that can implement the above-mentioned method.

[0079] Based on the above scheme, firstly, the audio data stream in the live video information obtained in real time is converted into a text sequence, and the initial sensitive text data group is extracted from the text sequence. Secondly, through the sensitive information screening network, the initial sensitive text data group obtained by the first screening is screened again to obtain each sensitive text data group. Finally, the audio data corresponding to the text data in each sensitive text data group in the audio data stream is muted, and then recombined with the obtained image data stream to obtain a new live data stream, and the new live data stream is sent to the client. Thereby, the protection effect of sensitive information during live broadcast is improved.

[0080] Optionally, based on a preset sensitive text extraction strategy, multiple groups of initial sensitive text data are extracted from each text data, including: in each text data, extracting each digital data in the text data through a digital feature extraction algorithm; storing continuous digital data in the same initial sensitive text data group to obtain multiple groups of initial sensitive text data.

[0081] In this embodiment, when the text data to be extracted is digital information, the terminal extracts each digital data in the text sequence through a digital feature extraction algorithm to obtain all the digital data in the text sequence; the terminal treats each adjacent digital data as a group of digital data according to the adjacent digital principle. Similarly, based on the above scheme, the terminal divides all digital data into multiple groups of digital data groups, and uses each group of digital data groups as the initial sensitive text data group. The digital feature extraction algorithm can be, but is not limited to, a rule matching method such as a regular expression. Digital information can be, but is not limited to, ID number, email address, amount, phone number, date, age, etc.

[0082] For example, if the text sequence is "Today is January 1, 2022", the terminal will extract a total of 6 digital data through the digital special diagnosis extraction algorithm. After obtaining each digital data, the terminal searches for adjacent digital data in the order of the text sequence, and divides the adjacent digital data into the same group, then a total of 3 groups of digital data are obtained, namely: "2022", the first "1", and the second "1".

[0083] In another embodiment, the terminal presets a special digital group category. After the terminal divides each digital data into each digital data group, it determines whether each digital data group meets the special digital group category. In the case that the digital data group meets the special digital group category, the terminal divides the digital data meeting the category and the adjacent text data in the text sequence into the same group to obtain a group of digital data. The digital group category can be, but is not limited to, ID number, email address, amount, mobile phone number.

[0084] The judgment condition codes for the above categories are:

[0085] Mobile phone number: [^(13[0-9]|14[5|7]|15[0|1|2|3|4|5|6|7|8|9]|18[0|1|2|3|5|6|7|8|9])\d{8}$];

[0086] ID number: [(^\d{15}$)|(^\d{18}$)|(^\d{17}(\d|X|x)$)];

[0087] Email address: ^\w+([-+.]\w+)*@\w+([-.]\w+)*\.\w+([-.]\w+)*$;

[0088] Amount (including decimal point): ^[0-9]+(.[0-9]+)?$.

[0089] For example, the text sequence is "The ID number is: 12345687987654321x, the email address is: 123456789@abc.cba, and the amount is: 66.66¥".

[0090] After the terminal extracts the text data through the digital feature extraction algorithm, it divides each adjacent digital data into the same group to obtain four groups of digital data, namely "12345687987654321", "123456789", the first "66", and the second "66". The terminal brings each group of digital data into the above-mentioned conditional judgment code to determine whether the four groups of digital data are special digital groups. Through the above-mentioned judgment operation, the first group of digital data is obtained to be a digital group belonging to the category of ID card number. Then, the terminal stores the adjacent English letters x after the digital data is sorted according to the text sequence into the group of digital data to obtain a special digital group "12345687987654321x". Similarly, through the above-mentioned judgment operation, three special digital groups are obtained, namely "12345687987654321x", 123456789@abc.cba, and "66.66¥".

[0091] Based on the above scheme, the digital data in the text sequence is extracted and divided through the digital feature extraction algorithm, thereby providing a data basis for further screening of sensitive text data.

[0092] Optionally, based on a preset sensitive text extraction strategy, multiple groups of initial sensitive text data are extracted from each text data, including: dividing each non-numeric data in the text sequence into multiple groups of non-numeric data through a division layer of a text extraction network; determining the category of each group of non-numeric data through a recognition layer of a text extraction network; and taking the non-numeric data corresponding to the preset category as initial sensitive text data to obtain multiple groups of initial sensitive text data.

[0093] In the present embodiment, when the text data to be extracted is non-numeric data, the terminal performs screening and partitioning operations on each non-numeric data in the text sequence through the partitioning layer of the text extraction network in the order of the text sequence, and screens out the non-numeric data corresponding to each sensitive information in each non-numeric data, and divides the adjacent non-numeric data into the same group to obtain each non-numeric data group. The specific partitioning operation will be described in detail later. The terminal presets the category information of each non-numeric data group, and first identifies the category information of each non-numeric data group through the recognition layer of the text extraction network, and divides each non-numeric data group into each category of non-numeric data groups according to the category, and selects the non-numeric data group of the preset category as the initial sensitive text data group in each category of non-numeric data groups, and obtains each initial sensitive text data group. The text extraction network can be, but is not limited to, a bidirectional long short-term memory network (Bi-long short term memory, BiLSTM). The partitioning layer is two long short-term neural networks (long short term memory, LSTM), and the recognition layer is a conditional random field constraint network (conditional random field, CRF).

[0094] The partitioning layer is used to filter and partition non-numeric data groups in the text sequence.

[0095] The recognition layer is used to recognize the categories of the non-digital data groups and select the non-digital data groups of the preset categories from each category.

[0096] For example, the text sequence is "Zhang San bought 50,000 financial products at Bank A yesterday...", and the terminal divides the non-numeric data into two groups through the division layer, namely "Zhang San" and "Bank A". The terminal obtains the category of "Zhang San" as name and the category of "Bank A" as company name through the recognition layer. When the terminal finally gives up the category of name, the terminal finally obtains the initial sensitive text data group of "Zhang San" through the text extraction network.

[0097] Based on the above solution, the terminal extracts non-numeric data from the text sequence through the text extraction network, thereby providing a data basis for further screening of sensitive text data.

[0098] Optionally, each non-numeric data in the text sequence is divided into multiple non-numeric data groups through the partitioning layer of the text extraction network, including: for each non-numeric data, judging whether the correlation between the non-numeric data and each non-numeric data adjacent to the non-numeric data reaches a preset correlation through the partitioning layer of the text extraction network; when the correlation between the non-numeric data and each non-numeric data adjacent to the non-numeric data is greater than the preset correlation, storing the non-numeric data and each non-numeric data adjacent to the non-numeric data into the same non-numeric data group to obtain multiple non-numeric data groups.

[0099] In this embodiment, the terminal presets the correlation between non-numeric data, and encodes each non-numeric data in the text sequence through the partitioning layer of the text extraction network. Then, for each non-numeric data in the text sequence, the terminal determines the correlation of the non-numeric data and the non-numeric data adjacent to the non-numeric data through the partitioning layer of the text extraction network, and determines whether the correlation of the non-numeric data and the non-numeric data adjacent to the non-numeric data is greater than the preset correlation. In the case where the correlation of the non-numeric data and the non-numeric data adjacent to the non-numeric data is greater than the preset correlation, the terminal divides the adjacent non-numeric data greater than the preset correlation into the same group to obtain each non-numeric data group.

[0100] Specifically, the partitioning layer includes a forward LSTM layer and a backward LSTM layer. The terminal judges each non-numeric data from front to back in the order of the text sequence through the forward LSTM layer, and whether the correlation of the non-numeric data adjacent to the non-numeric data is greater than the preset correlation, and screens out the non-numeric data greater than the preset correlation, and whether the correlation of the non-numeric data adjacent to the non-numeric data is greater than the preset correlation. The terminal judges each non-numeric data from back to front in the reverse order of the text sequence through the backward LSTM layer, and whether the correlation of the non-numeric data adjacent to the non-numeric data is greater than the preset correlation, and screens out the non-numeric data greater than the preset correlation, and whether the correlation of the non-numeric data adjacent to the non-numeric data is greater than the preset correlation. The terminal selects repeated non-numeric data from each non-numeric data obtained from the two screenings, and divides each non-numeric data adjacent to the repeated non-numeric data in the order of the text sequence into the same group to obtain each non-numeric data group.

[0101] For example, the text sequence is "Zhang San had nothing to do this afternoon, so he went to place A to buy groceries and had some steamed buns." The non-numeric data filtered by the forward LSTM at the terminal are "Zhang San," "Place A," and "Steamed buns." The non-numeric data filtered by the backward LSTM are "Zhang San," and "Place A." The terminal finally obtains two groups of non-numeric data, namely "Zhang San" and "Place A."

[0102] The terminal can simultaneously perform an extraction operation in which the initial sensitive text data in the text sequence is digital data, and an extraction operation in which the initial sensitive text data in the text sequence is non-digital data.

[0103] Based on the above scheme, the terminal screens and divides the non-numeric data of the text sequence into non-numeric data groups through the division layer of the text extraction network, thereby providing a data basis for the recognition layer of the subsequent text extraction network to perform recognition.

[0104] Optional, such as Figure 2 As shown, multiple groups of initial sensitive text data are input into the sensitive information screening network to obtain various sensitive text data groups, including:

[0105] Step S201: for each initial sensitive text data group, a confidence evaluation operation is performed on each initial sensitive text data in the initial sensitive text data group in a text sequence through a confidence evaluation network to obtain the confidence of the initial sensitive text data group.

[0106] Among them, the sensitive information screening network can be but is not limited to Bidirectional Encoder Representations from Transformers (BERT) based on transformers. The trained BERT network can determine the confidence of the labeled text data in the text sequence by identifying the contextual information in the text sequence.

[0107] In this embodiment, for each initial sensitive text data group, the terminal uses the sensitive information screening network in the text sequence to mark each initial sensitive text data in the initial sensitive text data group. The marking method can be, but is not limited to, inputting a vector representation of a special [CLS] identifier at the initial sensitive text data in the text sequence to mark the initial sensitive text data.

[0108] The terminal inputs the labeled text sequence into the confidence evaluation network (ie, the BERT network), and determines the confidence of each initial sensitive text data labeled in the text sequence through the confidence evaluation network.

[0109] Specifically, the terminal inputs the labeled text sequence into the BERT network, BERT recognizes the [CLS] identifier in the text sequence, and outputs the confidence level of each labeled initial sensitive text data in the range of 0 to 1 through BERT's Sigmoid classifier.

[0110] For example, the text sequence is "The above case of Zhang San being defrauded of 50,000 reminds us again that we should install the Anti-Fraud Center app and remember its phone number 96110", and the initial sensitive text data group is "Zhang San", "Anti-Fraud Center app", "96110". The terminal uses the BERT network to mark the initial sensitive text data "Zhang San" of the first group of initial sensitive data, and obtains the confidence of the text sequence marking the first group through the BERT network, that is, the confidence of "Zhang San" is 0.8. Similarly, through the above steps, the confidence of the text sequence marking the second group is obtained, that is, the confidence of "Anti-Fraud Center app" is 0.3; the confidence of the text sequence marking the third group, that is, the confidence of "96110" is 0.4.

[0111] Step S202: when the confidence level is greater than a preset confidence threshold, the initial sensitive text data group greater than the preset confidence threshold is used as a sensitive text data group to obtain various sensitive text data groups.

[0112] In this embodiment, the terminal presets a confidence threshold, and the terminal determines whether the confidence of all initial sensitive text data groups obtained through step S202 is greater than the preset confidence threshold, and screens out the initial sensitive text data groups greater than the preset confidence threshold as the sensitive text data groups.

[0113] For example, the text sequence is "The case of Zhang San being defrauded of 50,000 yuan reminds us again to install the anti-fraud center app and remember its phone number 96110", and the initial sensitive text data group is "Zhang San", "Anti-Fraud Center App", "96110". The confidence of "Zhang San" is 0.8, the confidence of "Anti-Fraud Center App" is 0.3, and the confidence of "96110" is 0.4. The terminal presets the confidence threshold to 0.5, so in the above-mentioned initial sensitive text data groups, the terminal "Zhang San" is used as the sensitive text data group.

[0114] Based on the above scheme, the extracted initial sensitive text data group is screened through a sensitive information screening network to obtain a sensitive text data group. This method can further improve the accuracy of the obtained sensitive text data group.

[0115] Optionally, the target audio data corresponding to the sensitive text data in each sensitive text data group in the audio data stream is muted to obtain a new audio data stream, including: in each audio data, screening the target audio data corresponding to each sensitive text data in the sensitive text data group; replacing the target audio data with preset audio data to obtain a new audio data stream.

[0116] In this embodiment, the terminal presets audio data, summarizes the sensitive text data in all sensitive text data groups, and extracts the start timestamp and end timestamp of the audio data of each pair of sensitive text data. The terminal replaces the audio data between each pair of start timestamps and end timestamps with preset audio data to obtain a new audio data stream. The preset audio data can be, but is not limited to, single audio data, messy audio data, no audio data, and other audio data that can meet the silencing requirements.

[0117] For example, the text sequence is "Zhang San's phone number is: 123456789, and his home address is: abcdefg." The sensitive text data corresponding to the sensitive text data group in the text sequence are "Zhang San", "123456789", and "abcdefg". The terminal replaces the audio data stream between the start timestamp and the end timestamp of the audio data corresponding to each sensitive text data with the preset audio data "beep", and obtains the text sequence corresponding to the new audio data stream as "Beep~'s phone number is: Beep~~~~~~, and his home address is: Beep~~~~~."

[0118] Based on the above scheme, the sensitive information of the audio data stream in the live data stream is replaced to obtain a new audio data stream, thereby improving the protection effect of sensitive information during live broadcast.

[0119] This application also provides an example of live data stream processing, such as Figure 3 As shown, the specific processing process includes the following steps:

[0120] Step S301, during the live broadcast process, obtain the live broadcast data stream.

[0121] Step S302: determining the text sequence corresponding to the audio data stream and the audio data corresponding to each text data in the text sequence in the audio data stream through a text conversion network.

[0122] Step S303: extract each digital data in each text data by using a digital feature extraction algorithm.

[0123] Step S304, storing the continuous digital data into the same initial sensitive text data group to obtain multiple groups of initial sensitive text data.

[0124] Step S305 , for each non-numeric data, judging whether the correlation between the non-numeric data and each non-numeric data adjacent to the non-numeric data reaches a preset correlation through the partitioning layer of the text extraction network.

[0125] Step S306, when the correlation between the non-digital data and each adjacent non-digital data is greater than a preset correlation, the non-digital data and each adjacent non-digital data are stored in the same non-digital data group to obtain multiple non-digital data groups.

[0126] Step S307, determining the category of each group of non-numeric data through the recognition layer of the text extraction network.

[0127] Step S308: Use the non-numeric data corresponding to the preset category as initial sensitive text data to obtain multiple groups of initial sensitive text data.

[0128] Step S309: for each initial sensitive text data group, a confidence evaluation operation is performed on each initial sensitive text data in the initial sensitive text data group in the text sequence through a confidence evaluation network to obtain the confidence of the initial sensitive text data group.

[0129] Step S310, when the confidence is greater than a preset confidence threshold, the initial sensitive text data group greater than the preset confidence threshold is used as a sensitive text data group to obtain various sensitive text data groups.

[0130] Step S311, screening target audio data corresponding to each sensitive text data in the sensitive text data group from among each audio data.

[0131] Step S312: Replace the target audio data with the preset audio data to obtain a new audio data stream.

[0132] Step S313: determine a new live data stream according to the new audio data stream and the image data stream, and send the new live data stream to the client.

[0133] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0134] Based on the same inventive concept, the embodiment of the present application also provides a live data stream processing device for implementing the live data stream processing method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more live data stream processing device embodiments provided below can refer to the limitations of the live data stream processing method above, and will not be repeated here.

[0135] In one embodiment, Figure 4 As shown, a live data stream processing device is provided, including: an acquisition module 410, an extraction module 420, a screening module 430, a muting module 440 and a sending module 450, wherein:

[0136] The acquisition module 410 is used to acquire a live broadcast data stream during the live broadcast process; the live broadcast data stream includes an image data stream and an audio data stream;

[0137] An extraction module 420 is used to determine, through a text conversion network, a text sequence corresponding to the audio data stream and audio data corresponding to each text data in the text sequence in the audio data stream;

[0138] A screening module 430 is used to extract multiple groups of initial sensitive text data from each of the text data based on a preset sensitive text extraction strategy, and input the multiple groups of initial sensitive text data into a sensitive information screening network to obtain each sensitive text data group;

[0139] A silencing module 440 is used to perform silencing processing on the target audio data corresponding to the sensitive text data in each of the sensitive text data groups in the audio data stream to obtain a new audio data stream;

[0140] The sending module 450 is used to determine a new live data stream according to the new audio data stream and the image data stream, and send the new live data stream to the client.

[0141] Optionally, the screening module 430 is specifically used to:

[0142] In each of the text data, extract each digital data in the text data by using a digital feature extraction algorithm;

[0143] The continuous digital data are stored in the same initial sensitive text data group to obtain multiple groups of the initial sensitive text data.

[0144] Optionally, the screening module 430 is specifically used to:

[0145] By using a partitioning layer of a text extraction network, each non-numeric data in the text sequence is divided into a plurality of groups of non-numeric data;

[0146] Determining the category of each group of the non-numeric data through the recognition layer of the text extraction network;

[0147] The non-numeric data corresponding to the preset category is used as the initial sensitive text data to obtain multiple groups of initial sensitive text data.

[0148] Optionally, the screening module 430 is specifically used to:

[0149] For each non-numeric data, judging whether the correlation between the non-numeric data and each non-numeric data adjacent to the non-numeric data reaches a preset correlation through the partitioning layer of the text extraction network;

[0150] When the correlation between the non-digital data and each non-digital data adjacent to the non-digital data is greater than a preset correlation, the non-digital data and each non-digital data adjacent to the non-digital data are stored in the same non-digital data group to obtain multiple non-digital data groups.

[0151] Optionally, the screening module 430 is specifically used to:

[0152] For each initial sensitive text data group, a confidence evaluation operation is performed on each initial sensitive text data in the initial sensitive text data group in the text sequence through a confidence evaluation network to obtain the confidence of the initial sensitive text data group;

[0153] In the case that the confidence is greater than a preset confidence threshold, the initial sensitive text data group greater than the preset confidence threshold is used as a sensitive text data group to obtain each of the sensitive text data groups.

[0154] Optionally, the muffler module 440 is specifically used for:

[0155] In each of the audio data, select the target audio data corresponding to each of the sensitive text data in the sensitive text data group;

[0156] The target audio data is replaced with preset audio data to obtain a new audio data stream.

[0157] Each module in the live data stream processing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0158] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a live data stream processing method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0159] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0160] In an embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0161] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0162] In an embodiment, a computer program product is provided, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0163] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0164] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0165] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for processing live data stream, It is characterized in that The method comprises: During the live broadcast process, a live broadcast data stream is obtained; the live broadcast data stream includes an image data stream and an audio data stream; Determine, through a text conversion network, a text sequence corresponding to the audio data stream, and audio data corresponding to each text data in the text sequence in the audio data stream; For each non-numeric data in the text sequence, through the forward LSTM layer of the partitioning layer of the text extraction network and the backward LSTM layer of the partitioning layer of the text extraction network, respectively determine whether the correlation between the non-numeric data and each non-numeric data adjacent to the non-numeric data reaches a preset correlation; When the correlation between the non-numeric data determined by the forward LSTM layer and each non-numeric data adjacent to the non-numeric data is greater than a preset correlation, and the correlation between the non-numeric data determined by the backward LSTM layer and each non-numeric data adjacent to the non-numeric data is greater than a preset correlation, the non-numeric data and each non-numeric data adjacent to the non-numeric data are stored in the same non-numeric data group to obtain multiple non-numeric data groups; Screening initial sensitive text data from each group of the non-digital data through the recognition layer of the text extraction network; Based on a preset sensitive text extraction strategy, extract multiple groups of initial sensitive text data from each of the text data; For each initial sensitive text data group, a confidence evaluation operation is performed on each initial sensitive text data in the initial sensitive text data group in the text sequence through a confidence evaluation network to obtain the confidence of the initial sensitive text data group; In the case where the confidence is greater than a preset confidence threshold, taking the initial sensitive text data group greater than the preset confidence threshold as the sensitive text data group, to obtain each of the sensitive text data groups; Silencing the target audio data corresponding to the sensitive text data in each of the sensitive text data groups in the audio data stream to obtain a new audio data stream; A new live data stream is determined according to the new audio data stream and the image data stream, and the new live data stream is sent to the client.

2. The method according to claim 1, It is characterized in that The step of extracting multiple groups of initial sensitive text data from each of the text data based on a preset sensitive text extraction strategy includes: In each of the text data, extract each digital data in the text data by using a digital feature extraction algorithm; The continuous digital data are stored in the same initial sensitive text data group to obtain multiple groups of the initial sensitive text data.

3. The method according to claim 1, It is characterized in that The step of screening the initial sensitive text data in each group of the non-digital data by the recognition layer of the text extraction network includes: Determining the category of each group of the non-numeric data through the recognition layer of the text extraction network; The non-numeric data corresponding to the preset category is used as the initial sensitive text data to obtain multiple groups of initial sensitive text data.

4. The method according to claim 1, It is characterized in that The step of muting the target audio data corresponding to the sensitive text data in each of the sensitive text data groups in the audio data stream to obtain a new audio data stream includes: In each of the audio data, select the target audio data corresponding to each of the sensitive text data in the sensitive text data group; The target audio data is replaced with preset audio data to obtain a new audio data stream.

5. A live data stream processing device, It is characterized in that The device comprises: An acquisition module is used to acquire a live broadcast data stream during the live broadcast process; the live broadcast data stream includes an image data stream and an audio data stream; An extraction module, used to determine, through a text conversion network, a text sequence corresponding to the audio data stream, and audio data corresponding to each text data in the text sequence in the audio data stream; The screening module is used for judging, for each non-numeric data in the text sequence, whether the correlation between the non-numeric data and each non-numeric data adjacent to the non-numeric data reaches a preset correlation through the forward LSTM layer of the partitioning layer of the text extraction network and the backward LSTM layer of the partitioning layer of the text extraction network; when the correlation between the non-numeric data judged by the forward LSTM layer and each non-numeric data adjacent to the non-numeric data is greater than the preset correlation, and the correlation between the non-numeric data judged by the backward LSTM layer and each non-numeric data adjacent to the non-numeric data is greater than the preset correlation, the non-numeric data is filtered out. The method further comprises: storing the non-numeric data, the non-numeric data adjacent to the non-numeric data, into the same non-numeric data group to obtain multiple non-numeric data groups; screening the initial sensitive text data in each group of the non-numeric data through the recognition layer of the text extraction network; for each initial sensitive text data group, performing a confidence evaluation operation on each initial sensitive text data in the initial sensitive text data group in the text sequence through the confidence evaluation network to obtain the confidence of the initial sensitive text data group; when the confidence is greater than a preset confidence threshold, taking the initial sensitive text data group greater than the preset confidence threshold as the sensitive text data group to obtain each of the sensitive text data groups; a silencing module, configured to perform silencing processing on target audio data corresponding to the sensitive text data in each of the sensitive text data groups in the audio data stream to obtain a new audio data stream; The sending module is used to determine a new live data stream according to the new audio data stream and the image data stream, and send the new live data stream to the client.

6. The device according to claim 5, It is characterized in that The screening module is specifically used for: In each of the text data, extract each digital data in the text data by using a digital feature extraction algorithm; The continuous digital data are stored in the same initial sensitive text data group to obtain multiple groups of the initial sensitive text data.

7. The device according to claim 5, It is characterized in that The screening module is specifically used for: Determining the category of each group of the non-numeric data through the recognition layer of the text extraction network; The non-numeric data corresponding to the preset category is used as the initial sensitive text data to obtain multiple groups of initial sensitive text data.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • A method and apparatus for detecting sensitive information

    CN110163013A

  • Live stream review intervention method and device, storage medium and equipment

    CN114339292A