A live video processing method, device and equipment

By using sensitive area detection and attention adjustment models, the problems of privacy leakage and attention distraction during live streaming are solved, and the automated processing of privacy protection and attention guidance is realized, thereby improving the efficiency and effectiveness of live streaming.

CN114428971BActive Publication Date: 2026-02-13ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210062129.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2026-02-13
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

During live video streaming, the streamer's privacy information and things unrelated to or competing with the products being sold may be leaked, affecting sales performance, and viewers' attention may be difficult to redirect in a timely manner.

Method used

A pre-trained sensitive area detection model is used to identify and de-identify sensitive areas in live videos. A smart contract is generated through a blockchain system to protect privacy. An attention detection model is used to adjust user attention and output adjustment strategies to guide user attention to preset areas.

Benefits of technology

It effectively protects the privacy of live streamers, improves the live streaming effect, enhances the attractiveness of live streaming content for users, reduces human intervention, and achieves automation and real-time adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114428971B_ABST
    Figure CN114428971B_ABST
Patent Text Reader

Abstract

The embodiment of the specification discloses a live video processing method, device and equipment, the method comprises: obtaining live video data stream to be displayed; inputting the live video data stream into a pre-trained sensitive area detection model to identify the sensitive area contained in the video corresponding to the live video data stream, based on the sensitive area information contained in the video, the sensitive information in the sensitive area contained in the video is desensitized to obtain the desensitized live video data stream, and the desensitized live video data stream is displayed, and the attention information of the user is detected; if the attention information of the user indicates that the attention of the user is not in the preset area of the video corresponding to the live video data stream, the adjustment strategy of the attention of the user is determined based on the attention information of the user and the desensitized live video data stream, and the adjustment strategy of the attention of the user is output to the live side corresponding to the live video data stream.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present document relates to the technical field of computers, and particularly relates to a live video processing method, device and equipment. BACKGROUND

[0002] It has become one of the current popular marketing methods to recommend and sell one or more commodities through video live streaming. Through the introduction and explanation of one or more commodities by a live streaming party (such as an anchor) in a network live streaming room opened by the live streaming party, a watching party can purchase the commodities through a shopping link provided by the network live streaming room. The above-mentioned sales method can remove many intermediate links in the sales process, thereby providing a more favorable price for the purchasing party and allowing a commodity producing party to remove the pressure of warehousing. However, in the process of video live streaming, the private information of the live streaming party and things unrelated to or in competition with the commodities to be sold may appear in the network live streaming room, thereby causing the leakage of private information and affecting the sales effect of the commodities to be sold. Therefore, a technical solution capable of more effectively protecting the privacy of the live streaming party and timely adjusting the attention of the watching party is needed. SUMMARY

[0003] The purpose of the embodiments of the present specification is to provide a technical solution capable of more effectively protecting the privacy of the live streaming party and timely adjusting the attention of the watching party.

[0004] In order to achieve the above-mentioned technical solution, the embodiments of the present specification are implemented as follows:

[0005] The live video processing method provided by the embodiments of the present specification comprises the following steps: obtaining live video data stream to be displayed. The live video data stream is input into a pre-trained sensitive area detection model to identify sensitive areas contained in a video corresponding to the live video data stream, and sensitive area information contained in the video is obtained. Based on the sensitive area information contained in the video, sensitive information in the sensitive areas contained in the video is desensitized, and desensitized live video data stream is obtained. The desensitized live video data stream is displayed, and the attention information of a user in the process of displaying the desensitized live video data stream is detected. If the attention information of the user indicates that the attention of the user is not in a preset area of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, an adjustment strategy of the attention of the user is determined, and the adjustment strategy of the attention of the user is output to a live streaming party corresponding to the live video data stream. The adjustment strategy of the attention of the user is used to instruct the live streaming party to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream.

[0006] The embodiment of the present specification provides a live video processing method, applied to a blockchain system, the method comprises: obtaining rule information for processing a live video, generating a corresponding first smart contract based on the rule information for processing the live video, and deploying the first smart contract into the blockchain system. Based on the first smart contract, obtain the live video data stream to be displayed. Based on the first smart contract, input the live video data stream into a pre-trained sensitive area detection model to identify the sensitive area contained in the video corresponding to the live video data stream, and obtain the sensitive area information contained in the video. Based on the first smart contract and the sensitive area information contained in the video, the sensitive information in the sensitive area contained in the video is desensitized, and the desensitized live video data stream is obtained. Based on the first smart contract, the desensitized live video data stream is displayed, and the attention information of the user in the process of displaying the desensitized live video data stream is detected. If the user's attention information indicates that the user's attention is not in the preset area of the video corresponding to the live video data stream, based on the first smart contract, the user's attention information and the desensitized live video data stream, the adjustment strategy of the user's attention is determined, and the adjustment strategy of the user's attention is output to the live party corresponding to the live video data stream. The adjustment strategy of the user's attention is used to instruct the live party to guide the user's attention to be in the preset area of the video corresponding to the live video data stream.

[0007] The embodiment of the present specification provides a live video processing device, the device comprises: a data acquisition module, acquiring a live video data stream to be displayed. A sensitive area identification module inputs the live video data stream into a pre-trained sensitive area detection model to identify the sensitive area contained in the video corresponding to the live video data stream, and obtains the sensitive area information contained in the video. A desensitization module, based on the sensitive area information contained in the video, desensitizes the sensitive information in the sensitive area contained in the video, and obtains the desensitized live video data stream. An attention detection module displays the desensitized live video data stream, and detects the attention information of the user in the process of displaying the desensitized live video data stream. An attention adjustment module, if the user's attention information indicates that the user's attention is not in the preset area of the video corresponding to the live video data stream, based on the user's attention information and the desensitized live video data stream, the adjustment strategy of the user's attention is determined, and the adjustment strategy of the user's attention is output to the live party corresponding to the live video data stream. The adjustment strategy of the user's attention is used to instruct the live party to guide the user's attention to be in the preset area of the video corresponding to the live video data stream.

[0008] The embodiment of the specification provides a live video processing device. The device is a device in a blockchain system. The device comprises: a contract deployment module, which acquires rule information for processing a live video, generates a corresponding first smart contract based on the rule information for processing the live video, and deploys the first smart contract into the blockchain system. A data acquisition module acquires a live video data stream to be displayed based on the first smart contract. A sensitive area identification module inputs the live video data stream into a pre-trained sensitive area detection model based on the first smart contract, identifies sensitive areas contained in a video corresponding to the live video data stream, and obtains sensitive area information contained in the video. A desensitization module desensitizes sensitive information in the sensitive areas contained in the video based on the first smart contract and the sensitive area information contained in the video, and obtains a desensitized live video data stream. An attention detection module displays the desensitized live video data stream based on the first smart contract, and detects attention information of a user in a process of displaying the desensitized live video data stream. An attention adjustment module determines an adjustment strategy of the user's attention based on the first smart contract, the user's attention information, and the desensitized live video data stream if the user's attention information indicates that the user's attention is not in a preset area of the video corresponding to the live video data stream, and outputs the adjustment strategy of the user's attention to a live streaming party corresponding to the live video data stream. The adjustment strategy of the user's attention is used to instruct the live streaming party to guide the user's attention to be in the preset area of the video corresponding to the live video data stream.

[0009] An embodiment of the specification provides a live video processing device, comprising: a processor; and a memory arranged to store computer executable instructions that, when executed, cause the processor to: acquire a live video data stream to be displayed. input the live video data stream into a pre-trained sensitive region detection model to identify sensitive regions contained in a video corresponding to the live video data stream, to obtain sensitive region information contained in the video. based on the sensitive region information contained in the video, perform desensitization processing on sensitive information in the sensitive regions contained in the video, to obtain a desensitized live video data stream. display the desensitized live video data stream, and detect attention information of a user in the process of displaying the desensitized live video data stream. if the attention information of the user indicates that the attention of the user is not in a preset region of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, determine an adjustment strategy for the attention of the user, and output the adjustment strategy for the attention of the user to a live streaming party corresponding to the live video data stream, the adjustment strategy for the attention of the user being used to instruct the live streaming party to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream.

[0010] An embodiment of the present specification provides a live video processing device. The device is a device in a blockchain system. The live video processing device comprises a processor and a memory arranged to store computer executable instructions that, when executed, cause the processor to: obtain rule information for processing a live video, generate a corresponding first smart contract based on the rule information for processing the live video, and deploy the first smart contract to the blockchain system. Obtain a live video data stream to be displayed based on the first smart contract. Input the live video data stream into a pre-trained sensitive region detection model based on the first smart contract to identify sensitive regions contained in a video corresponding to the live video data stream, and obtain sensitive region information contained in the video. Desensitize sensitive information in the sensitive regions contained in the video based on the first smart contract and the sensitive region information contained in the video, and obtain a desensitized live video data stream. Display the desensitized live video data stream based on the first smart contract, and detect attention information of a user in the process of displaying the desensitized live video data stream. If the attention information of the user indicates that the attention of the user is not in a preset region of the video corresponding to the live video data stream, determine an adjustment strategy for the attention of the user based on the first smart contract, the attention information of the user, and the desensitized live video data stream, and output the adjustment strategy for the attention of the user to a live streaming party corresponding to the live video data stream. The adjustment strategy for the attention of the user is used to instruct the live streaming party to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream.

[0011] The embodiment of the present specification further provides a storage medium, wherein the storage medium is used to store computer executable instructions, and the executable instructions implement the following processes when executed: obtaining a live video data stream to be displayed. Inputting the live video data stream into a pre-trained sensitive area detection model to identify sensitive areas contained in a video corresponding to the live video data stream, to obtain sensitive area information contained in the video. Based on the sensitive area information contained in the video, performing desensitization processing on sensitive information in the sensitive areas contained in the video to obtain a desensitized live video data stream. Displaying the desensitized live video data stream, and detecting attention information of a user in the process of displaying the desensitized live video data stream. If the attention information of the user indicates that the attention of the user is not in a preset area of the video corresponding to the live video data stream, determining an adjustment strategy of the attention of the user based on the attention information of the user and the desensitized live video data stream, and outputting the adjustment strategy of the attention of the user to a live streaming party corresponding to the live video data stream, the adjustment strategy of the attention of the user being used to instruct the live streaming party to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream.

[0012] The embodiment of the present specification further provides a storage medium, wherein the storage medium is used to store computer executable instructions, and the executable instructions implement the following processes when executed: obtaining rule information for processing a live video, generating a corresponding first smart contract based on the rule information for processing the live video, and deploying the first smart contract into the blockchain system. Obtaining a live video data stream to be displayed based on the first smart contract. Inputting the live video data stream into a pre-trained sensitive area detection model based on the first smart contract to identify sensitive areas contained in a video corresponding to the live video data stream, to obtain sensitive area information contained in the video. Based on the first smart contract and the sensitive area information contained in the video, performing desensitization processing on sensitive information in the sensitive areas contained in the video to obtain a desensitized live video data stream. Based on the first smart contract, displaying the desensitized live video data stream, and detecting attention information of a user in the process of displaying the desensitized live video data stream. If the attention information of the user indicates that the attention of the user is not in a preset area of the video corresponding to the live video data stream, determining an adjustment strategy of the attention of the user based on the first smart contract, the attention information of the user and the desensitized live video data stream, and outputting the adjustment strategy of the attention of the user to a live streaming party corresponding to the live video data stream, the adjustment strategy of the attention of the user being used to instruct the live streaming party to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is an embodiment of a live video processing method described in this specification;

[0015] Figure 2 This is a schematic diagram of a live video interface as described in this specification;

[0016] Figure 3A This is a schematic diagram of another live video interface in this manual;

[0017] Figure 3B This is a schematic diagram of another live video interface in this manual;

[0018] Figure 4 This is another embodiment of a live video processing method described in this specification;

[0019] Figure 5A This is yet another embodiment of a live video processing method described in this specification;

[0020] Figure 5B This is a schematic diagram illustrating a live video processing procedure as described in this specification.

[0021] Figure 6 This specification provides an embodiment of a live video processing device.

[0022] Figure 7 This is another embodiment of a live video processing device described in this specification;

[0023] Figure 8 This is an embodiment of a live video processing device described in this specification. Detailed Implementation

[0024] This specification provides an embodiment of a method, apparatus, and device for processing live video.

[0025] In order for those skilled in the technical field to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described in the specification below in conjunction with the drawings in the specification. Obviously, the described embodiments are only part of the embodiments of the specification, not all. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative labor should be within the scope of protection of the specification.

[0026] Embodiment one

[0027] As Figure 1 shown, the embodiments of the specification provide a live video processing method, the execution subject of the method can be a terminal device or a server, wherein the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, etc., or a terminal device such as a personal computer, etc. The server can be an independent server, or a server cluster composed of multiple servers, etc., which can be a background server of a financial service or an online shopping service, etc., or a background server of an application program, etc. The method can be applied to live video processing and other related scenarios, and the method can specifically include the following steps:

[0028] In step S102, the live video data stream to be displayed is obtained.

[0029] Among them, the live video data stream can be a data stream in which a certain live party provides certain information to the viewer through video live streaming, for example, the live video data stream can be a data stream in which the live party recommends a certain commodity to the viewer through video live streaming, or it can also be a data stream in which the live party provides a certain online performance to the viewer through video live streaming, or it can also be a data stream in which the live party displays the daily life of the live party to the viewer through video live streaming, etc., which can be set according to actual conditions, and the embodiments of the specification do not limit it.

[0030] In practice, recommending and selling one or more products through live video streaming has become a popular marketing method. The streamer (e.g., a host) introduces and explains products in their online live streaming room, and viewers can purchase them through shopping links provided in the live stream. This sales method eliminates many intermediaries in the sales process, offering buyers more favorable prices and relieving manufacturers of warehousing pressure. However, during live video streaming, the streamer's private information, as well as items unrelated to or competing with the products being sold, may appear in the live stream, leading to privacy leaks and affecting sales performance. For example, if a streamer is broadcasting from their studio, viewers can infer or learn about the streamer or studio from the view outside the window or documents on the desk, resulting in privacy leaks. Similarly, if a streamer or their assistant is wearing clothing or accessories from another brand while introducing shoes, it distracts viewers and negatively impacts sales. Therefore, there is a need for a technical solution that can more effectively protect the privacy of the live streamer and adjust the viewer's attention in a timely manner. This specification provides an optional implementation method, which may include the following:

[0031] There are several ways to obtain the live video data stream to be displayed. For example, the live streamer's terminal device can have a live video streaming application installed. When the live streamer needs to conduct a live video stream, they can launch this application. This application can be configured with a trigger mechanism for opening or entering the live stream room, for example... Figure 2 As shown, the application can include buttons or hyperlinks for opening or entering a live streaming room. The streamer can trigger the process of opening or entering the live streaming room based on the aforementioned triggering mechanism. The application can execute the steps for opening or entering the live streaming room. Then, the application can call the camera component of the terminal device to collect relevant information such as the streamer and the environment within a specified range. This collected information can be displayed on the terminal device's display component through the application. Before displaying the collected information on the terminal device's display component, the live video data stream to be displayed can be acquired. Besides the above method, various other methods can be used, depending on the actual situation. This specification does not limit the specific methods used in this embodiment.

[0032] In step S104, the live video data stream is input into the pre-trained sensitive region detection model to identify the sensitive region contained in the video corresponding to the live video data stream, to obtain the sensitive region information contained in the video.

[0033] The sensitive region detection model can be a model for detecting sensitive regions contained in a live video. The sensitive region detection model can be constructed by various algorithms or network models, for example, the sensitive region detection model can be constructed by a classification algorithm, an image recognition algorithm, etc. The sensitive region can be a region containing sensitive information (such as the privacy information (such as the number of a certain certificate, a mobile phone number, etc.) of the live party, the relevant information (such as a certain clause in a contract of the live party, etc.) in a certain file of the live party, or certain sensitive information (such as the scenery outside the window, a certain decoration on the clothes of the live party, clothes, etc.) in the current live environment, etc. In addition to the above-mentioned sensitive information region, it can also include various sensitive information regions. The specific implementation can be determined according to actual conditions, and the present application does not limit the specific implementation.

[0034] In implementation, in order to improve the detection efficiency and accuracy of sensitive information, the sensitive region detection model can be constructed based on a pre-set algorithm or network model, and a plurality of training samples can be obtained. The sensitive region detection model can be trained by the training samples. Finally, the trained sensitive region detection model is obtained. The trained sensitive region detection model can detect the region of sensitive information contained in the video. It can also identify the content and type of sensitive information in the sensitive region, such as clothes, shoes or accessories, brand labels on clothes, brand labels on shoes, characters engraved in accessories, etc. In addition to the above-mentioned functions, the sensitive region detection model can also include various other functions. The specific implementation can be determined according to actual conditions, and the present application does not limit the specific implementation. After obtaining the live video data stream in the above-mentioned manner, the live video data stream can be input into the trained sensitive region detection model. The sensitive region detection model can identify the sensitive region contained in the video corresponding to the live video data stream, so as to detect the sensitive region contained in the video corresponding to the live video data stream. The sensitive region can be displayed in the form of coordinates or highlighted in the above-mentioned video (for example, the sensitive region can be circled with a specified color circle or rectangle, etc.). The specific implementation can be determined according to actual conditions, and the present application does not limit the specific implementation. As mentioned above, in addition to determining the sensitive region, the content and type of sensitive information in the sensitive region can also be identified, to obtain the content and type of sensitive information in each sensitive region.

[0035] In step S106, based on the sensitive region information contained in the video, the sensitive information in the sensitive region contained in the video is desensitized to obtain the desensitized live video data stream.

[0036] In implementation, the desensitization mechanism of sensitive information can be set in advance, which can include multiple mechanisms, for example, using a specified image to replace the image of the sensitive region where the sensitive information is located, or performing mosaic processing on the sensitive region where the sensitive information is located, etc. In addition, different desensitization mechanisms can be used based on different types of sensitive information. The specific setting can be determined according to actual conditions, and the embodiments of the present specification are not limited. After obtaining the sensitive region information contained in the video in the above manner, the sensitive region information can be analyzed to determine the sensitive region contained in the video and the content and type of the sensitive information contained in each sensitive region. The type of sensitive information contained in each sensitive region can be obtained. Then, the desensitization mechanism can be used to desensitize the corresponding sensitive information, so that the sensitive information in the sensitive region contained in the video is desensitized to obtain the desensitized live video data stream.

[0037] In step S108, the desensitized live video data stream is displayed, and the attention information of the user in the process of displaying the desensitized live video data stream is detected.

[0038] In implementation, after obtaining the desensitized live video data stream in the above manner, the current live video data stream no longer contains sensitive information. At this time, the desensitized live video data stream can be provided to the display component. The desensitized live video data stream can be displayed through the display component, and the viewer (i.e. the user) can watch the desensitized live video data stream without sensitive information. In addition, in order to improve the user's viewing efficiency of the live video and improve the attraction of the content of the live video to the user, the user's attention information can be obtained by detecting the region or range of the user's attention or gaze in the live video during the display of the desensitized live video data stream.

[0039] In step S110, if the attention information of the user indicates that the attention of the user is not in the preset region of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, the adjustment strategy of the attention of the user is determined, and the adjustment strategy of the attention of the user is output to the live side corresponding to the live video data stream. The adjustment strategy of the attention of the user is used to instruct the live side to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream.

[0040] The attention adjustment strategy can be a strategy of adjusting the user's attention so that the user's attention or focus (or gaze) is in a preset area (such as an area where a product introduced by the live streaming party is located or an area where the live streaming party is located, etc.). The attention adjustment strategy can be embodied in the form of text, for example, as shown in Figure 3A The live streaming party can adjust the user's attention through the text provided in the script content, or can be embodied in the form of a combination of multiple types of text, images, videos, and audio, for example, as shown in Figure 3B The live streaming party can adjust the user's attention through the text provided in the script content and the action provided in the video. The specific adjustment can be set according to actual conditions, and the embodiments of the present specification do not limit the adjustment.

[0041] In implementation, in order to improve the user's viewing efficiency of the live streaming video and the attraction of the content of the live streaming video to the user, the attention adjustment strategy in different scenes or situations can be set in advance. The set attention adjustment strategy can be set based on analysis of big data or based on expert experience, etc. If the attention information of the user obtained through the above method indicates that the user's attention is not in the preset area of the video corresponding to the live streaming video data stream, it indicates that the user is attracted by something other than the main key matter in the current live streaming video. At this time, based on the analysis of the attention information of the user and the desensitized live streaming video data stream, the current scene or the situation to which the user belongs can be obtained, and then the corresponding attention adjustment strategy of the user can be obtained. The attention adjustment strategy of the user can be output to the live streaming party corresponding to the live streaming video data stream. After receiving the attention adjustment strategy of the user, the live streaming party can apply the attention adjustment strategy of the user in the current live streaming video based on the mode recorded in the attention adjustment strategy of the user, so as to guide the user's attention to be in the preset area of the video corresponding to the live streaming video data stream again.

[0042] The embodiment of the present specification provides a live video processing method, obtaining live video data stream to be displayed, inputting the live video data stream into a pre-trained sensitive area detection model to identify the sensitive area contained in the video corresponding to the live video data stream, obtaining the sensitive area information contained in the video, then, based on the sensitive area information contained in the video, desensitizing the sensitive information in the sensitive area contained in the video to obtain the desensitized live video data stream, displaying the desensitized live video data stream, and detecting the attention information of the user in the process of displaying the desensitized live video data stream, if the attention information of the user indicates that the attention of the user is not in the preset area of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, determining the adjustment strategy of the user's attention, and outputting the adjustment strategy of the user's attention to the live side corresponding to the live video data stream, the adjustment strategy of the user's attention is used to instruct the live side to guide the user's attention to be in the preset area of the video corresponding to the live video data stream, in this way, the deep learning model (i.e. sensitive area detection model) is used to convert a large amount of manual work into an automatic, quasi-real-time standardized process, and the sensitive area detection model can automatically locate to private information / information irrelevant to the live object, and then process through the desensitization technology without manual intervention, and in terms of attention attraction, the user's attention can also be automatically detected, and the optimal adjustment strategy of the attention is automatically predicted and given to the live side, without relying on the experience of the live side, improving the live efficiency.

[0043] Embodiment two

[0044] As Figure 4 shown, the embodiment of the present specification provides a live video processing method, and the execution subject of the method can be a terminal device or a server, wherein the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, etc., or a terminal device such as a personal computer, etc. The server can be an independent server, or a server cluster composed of multiple servers, etc., and the server can be a background server of a financial service or a network shopping service, etc., or a background server of an application program, etc. The method can be applied to live video processing and other related scenes, and the method can specifically include the following steps:

[0045] In actual application, a corresponding model can be constructed by machine learning, and the sensitive area in the video corresponding to the live video data stream is detected through the model, which can specifically include the processing of the following steps S402 and S404.

[0046] In step S402, a first training sample is obtained, the first training sample is composed of historical live video data stream, and the video corresponding to the historical live video data stream contains a sensitive area.

[0047] In implementation, the live video data stream generated in the video live broadcast process of one or more different live broadcasters can be obtained in various different ways, for example, the live video data stream generated in the video live broadcast process of one or more different live broadcasters can be obtained by exchange such as purchase, or the live video data stream can be searched in the Internet or a specified local area network by a web crawler, etc. The live video data stream obtained by the above-mentioned way can be analyzed to determine the sensitive area contained therein. The sensitive area can include a sensitive area in the space environment where the live broadcaster is located (specifically, such as the scenery outside the window or the landscape of the window), a preset type of article of the live broadcaster or a local area of the article (specifically, such as the area where the text in a certain contract is located, the area where the information of the address field in a certain certificate is located, etc.), an area where the relevant information of the live object corresponding to the live video data stream is located (specifically, such as a brand or a non-brand product corresponding to the brand related to the live object), etc. The sensitive area contained in the video corresponding to each live video data stream can be labeled by a label labeling method. The labeled live video data stream can be used as a historical live video data stream, and the obtained historical live video data stream can be used as a first training sample.

[0048] In step S404, the sensitive area detection model is trained based on the first training sample and a first loss function based on mean square error, to obtain a trained sensitive area detection model. The sensitive area detection model is a model constructed based on a neural network algorithm, and the neural network algorithm includes a FasterRCNN algorithm.

[0049] In implementation, the algorithm for constructing the sensitive area detection model can be pre-set. The algorithm can include various algorithms, for example, a neural network algorithm, which can include various algorithms, for example, a convolutional neural network algorithm or a FasterRCNN algorithm, etc. The neural network algorithm in the present embodiment can be a FasterRCNN algorithm. The FasterRCNN algorithm can also include various implementable forms, for example, the first training sample obtained above can be input into a CNN for feature extraction, and a RPN is used to generate a proposal size or a proposal size (Proposals). Each first training sample generates a plurality of proposal sizes or proposal sizes. The proposal size or proposal size is mapped to the last convolutional feature map of the CNN. Each RoI generates a fixed size or fixed size feature map through a RoIPooling layer. The first loss function based on mean square error is used to jointly train the classification probability and the bounding box regression until the trained sensitive area detection model reaches the pre-set accuracy (for example, 95%, etc.). Finally, the trained sensitive area detection model can be obtained.

[0050] It should be noted that in actual application, the first training sample can be divided into two parts, one part is used for model training, and the other part can be used for testing the trained model. Specifically, the sample set composed of the first training sample can be divided into 80% and 20% parts, the 80% part can be used for model training, and the 20% part can be used for testing the trained model.

[0051] In actual application, a corresponding model can be constructed by machine learning, and the model is used to detect the attention of the user watching the live video. Specifically, the following steps S406 and S408 can be included.

[0052] In step S406, a second training sample carrying a user attention map is obtained, and the second training sample is composed of historical live video data streams.

[0053] In implementation, a plurality of different users can be recruited or invited to wear an eye tracker to watch video live. The eye tracker will record the user's attention map, so that the second training sample carrying the user's attention map can be obtained. In actual application, the second training sample can be divided into two parts, one part is used for model training, and the other part can be used for testing the trained model. Specifically, the sample set composed of the second training sample can be divided into 80% and 20% parts, the 80% part can be used for model training, and the 20% part can be used for testing the trained model.

[0054] In step S408, the attention detection model is trained based on the second training sample and the second loss function based on mean square error, and the trained attention detection model is obtained. The attention detection model is a model constructed based on image segmentation UNet algorithm.

[0055] In implementation, the input data of the attention detection model can be the second training sample carrying the user's attention map, and the output data can be the information of the user's attention area (which can be presented in the form of attention map). The structure of the attention detection model can adopt a network structure similar to that of saliency detection, for example, a model based on image segmentation UNet algorithm. The second training sample carrying the user's attention map and the attention detection model can be used to train the model on the second loss function based on mean square error, until the trained attention detection model reaches the pre-set accuracy (for example, 95%) on the test set. Finally, the trained attention detection model can be obtained.

[0056] In actual application, a corresponding model can be constructed by machine learning, and the model is used to adjust the attention of the user watching the live video. Specifically, the following steps S410 and S412 can be included.

[0057] In step S410, a third training sample carrying an adjustment strategy of user attention and a corresponding attention map is obtained, and the third training sample is composed of historical live video data streams.

[0058] In implementation, the manner of obtaining the third training sample can include multiple manners. An optional processing manner is provided below, which can include the following content: live video data streams of one or more different live parties for video live broadcast are collected, and the attention map of the user corresponding to each live video data stream can be predicted by using the above attention detection model. In addition, before and after the live party significantly adjusts the position (such as zooming out, zooming in, rotating, etc.) of the live object (such as goods, etc.) or uses a certain rhetoric, the change of the user's attention map can be compared. If the user's attention map has a significant change, the action, action type and rhetoric content of the live party, as well as the attention map corresponding to the video live content + rhetoric / action can be recorded. In addition, the above action and rhetoric can be classified, so as to divide the action and rhetoric into several categories. Then, {live video data stream + video live content + attention map corresponding to the rhetoric / action, action / rhetoric category} can be taken as the third training sample.

[0059] It should be noted that the third training sample can be divided into two parts, one part is used for model training, and the other part can be used for testing the trained model. Specifically, the sample set composed of the third training sample can be divided into 80% and 20% parts, the 80% part can be used for model training, and the 20% part can be used for testing the trained model.

[0060] In step S412, the attention adjustment model is trained based on the third training sample and a preset Softmax loss function, and a trained attention adjustment model is obtained. The attention adjustment model is a model constructed based on a residual network ResNet algorithm.

[0061] In implementation, the attention adjustment model can be constructed by various different algorithms, for example, the attention adjustment model can be constructed based on a neural network algorithm, the attention adjustment model can also be constructed based on a classification algorithm, and the attention adjustment model can also be constructed based on a residual network ResNet algorithm, etc. Taking the attention adjustment model constructed based on the residual network ResNet algorithm as an example, the residual network ResNet algorithm in it can be, for example, ResNet50, ResNet101, etc. Each third training sample obtained above can be input into the attention adjustment model to obtain a corresponding output result. Through the output result, a preset Softmax loss function is used to calculate a corresponding loss value. Based on the obtained loss value, the model parameters in the attention adjustment model are adjusted to obtain an updated attention adjustment model. Then the above processing process is repeated until the trained attention adjustment model reaches a pre-set accuracy (for example, 98%) on the test set. Finally, the trained attention adjustment model can be obtained.

[0062] After obtaining the models with different functions and purposes in the above manner, the above models can be applied to process the video live broadcast, which can specifically include the following steps S414 to step S426.

[0063] In step S414, live video data stream to be displayed is obtained.

[0064] In step S416, the live video data stream is input into the pre-trained sensitive area detection model to identify the sensitive area contained in the video corresponding to the live video data stream, and obtain the sensitive area information contained in the video.

[0065] The sensitive area information can include information of a sensitive area in a space environment where the live broadcast party is located, information of a preset type of object of the live broadcast party or a local area of the object, and information related to a live broadcast object corresponding to the live video data stream.

[0066] In step S418, the sensitive information in the sensitive area corresponding to the sensitive area information contained in the above video is identified based on the pre-trained information identification model, and the sensitive information in the sensitive area contained in the video is obtained.

[0067] In step S420, the sensitive information in the sensitive area contained in the above video is desensitized, and the desensitized live video data stream is obtained.

[0068] In implementation, the trained sensitive region detection model can be used to detect the content of the live video corresponding to the live video data stream to be displayed in real time, and after obtaining the sensitive region, the information recognition model can be used to identify the sensitive information in the sensitive region corresponding to the sensitive region information contained in the video, and obtain the sensitive information in the sensitive region contained in the video. Then, the sensitive information in the sensitive region can be desensitized by using desensitization operations such as digital stickers.

[0069] The information desensitization of the live object-independent thing can include: using the trained sensitive region detection model to detect the content of the live video corresponding to the live video data stream to be displayed in real time, and obtaining the sensitive region, and further obtaining the corresponding brand and commodity type and other related information through the information recognition model. If it is a product with the same brand as the product to be sold in the current live video, desensitization is not required, otherwise, the sensitive information in the sensitive region can be desensitized by using desensitization operations such as digital stickers.

[0070] In step S422, the desensitized live video data stream is displayed, and the desensitized live video data stream is input into the pre-trained attention detection model to obtain the attention information of the user during the display of the desensitized live video data stream.

[0071] In implementation, the trained attention detection model can be used to predict the user's attention map in real time during the display of the desensitized live video data stream, and obtain the user's attention information. If the high response area of the user's attention map is highly coincident with the area where the live object is located (i.e. the preset area) (such as the coincidence rate IoU exceeding 50% is considered to be highly coincident), it means that the user's attention is focused on the live object, and the user's attention adjustment is not required.

[0072] In step S424, if the user's attention information indicates that the user's attention is not in the preset area of the video corresponding to the live video data stream, the user's attention information and the desensitized live video data stream are input into the pre-trained attention adjustment model to obtain the adjustment strategy of the user's attention. The adjustment strategy of the user's attention is used to indicate that the live party guides the user's attention to be in the preset area of the video corresponding to the live video data stream.

[0073] The adjustment strategy of the attention includes a speech adjustment strategy and / or a behavior action adjustment strategy.

[0074] In implementation, if the high response area of the user's attention map is not in the area where the live object is located (i.e., the preset area), an attention map of the high response area in the area where the live object is located is created, the attention map and the desensitized live video data stream are input into the trained attention adjustment model, and the output is a speech adjustment strategy and / or a behavior action adjustment strategy. Then, according to the speech adjustment strategy and / or the behavior action adjustment strategy, the live party is prompted to perform the corresponding action and / or speech, so as to further manage the user's attention and attract it to the live object.

[0075] In step S426, the adjustment strategy of the user's attention is output to the live party corresponding to the live video data stream.

[0076] The embodiment of the present specification provides a processing method of live video, obtaining a live video data stream to be displayed, inputting the live video data stream into a pre-trained sensitive area detection model to identify a sensitive area contained in a video corresponding to the live video data stream, obtaining sensitive area information contained in the video, then, based on the sensitive area information contained in the video, desensitizing sensitive information in the sensitive area contained in the video to obtain a desensitized live video data stream, displaying the desensitized live video data stream, and detecting attention information of a user in the process of displaying the desensitized live video data stream. If the user's attention information indicates that the user's attention is not in a preset area of a video corresponding to the live video data stream, the adjustment strategy of the user's attention is determined based on the user's attention information and the desensitized live video data stream, and the adjustment strategy of the user's attention is output to the live party corresponding to the live video data stream. The adjustment strategy of the user's attention is used to instruct the live party to guide the user's attention to be in the preset area of the video corresponding to the live video data stream. In this way, the deep learning model (i.e., the sensitive area detection model) is used to convert a large amount of manual work into an automatic, quasi-real-time standardized process, and the sensitive area detection model can automatically locate private information / irrelevant information to the live object, and then process it through desensitization technology without human intervention. In terms of attention attraction, the user's attention can also be automatically detected, and the optimal adjustment strategy of the attention is automatically predicted and given to the live party, without relying on the experience of the live party, thereby improving the live efficiency.

[0077] Embodiment three

[0078] As Figure 5A and Figure 5BAs shown, the embodiment of the present specification provides a live video processing method, the execution subject of the method can be a blockchain system, the blockchain system can be composed of a terminal device or a server, etc., wherein the server can be a background server for a certain service (such as a transaction service or a financial service, etc.) or providing access to a certain transaction object, specifically, the server can be a server of a payment service, or a server related to a financial or instant messaging service, etc. The method can be applied to related scenes of live video processing, and the method can specifically include the following steps:

[0079] In step S502, the rule information for processing the live video is obtained, the corresponding first smart contract is generated based on the rule information for processing the live video, and the first smart contract is deployed to the blockchain system.

[0080] Among them, the first smart contract can be a computer protocol designed to disseminate, verify or execute contracts in an information-based way, the first smart contract allows trusted interaction without a third party, the above-mentioned interaction process is traceable and irreversible, and the first smart contract includes an agreement that contract participants can execute on the agreement that contract participants agree to the rights and obligations.

[0081] In implementation, in order to make the traceability of processing the live video better, a specified blockchain system can be created or joined, so that the live video can be processed based on the blockchain system. Specifically, the blockchain node can be installed with a corresponding application, the application can be provided with an input box and / or a selection box for the rule information for processing the live video, etc., and the corresponding information can be set in the input box and / or the selection box. Then, the blockchain system can receive the rule information for processing the live video. The blockchain system can generate a corresponding first smart contract based on the rule information for processing the live video, and can deploy the first smart contract to the blockchain system, so that the rule information for processing the live video and the corresponding first smart contract are stored in the blockchain system, other users cannot tamper with the rule information for processing the live video and the corresponding first smart contract, and the blockchain system processes the live video through the first smart contract.

[0082] In step S504, the live video data stream to be displayed is obtained based on the first smart contract.

[0083] In implementation, the first smart contract can be provided with related rule information for obtaining the live video data stream to be displayed, so that the above-mentioned corresponding processing can be realized based on the above-mentioned rule information in the first smart contract, and specific reference can be made to the above-mentioned related content, which will not be repeated here.

[0084] In step S506, based on the first smart contract, the live video data stream is input into the pre-trained sensitive region detection model to identify the sensitive region contained in the video corresponding to the live video data stream, and sensitive region information contained in the video is obtained.

[0085] In implementation, the first smart contract can be provided with related rule information of inputting the live video data stream into the pre-trained sensitive region detection model to identify the sensitive region contained in the video corresponding to the live video data stream. Thus, the above-mentioned corresponding processing can be realized based on the above-mentioned rule information in the first smart contract. For details, please refer to the above-mentioned related content, which will not be repeated here.

[0086] In step S508, based on the first smart contract and the sensitive region information contained in the video, the sensitive information in the sensitive region contained in the video is desensitized to obtain the desensitized live video data stream.

[0087] In implementation, the first smart contract can be provided with related rule information of desensitizing the sensitive information in the sensitive region contained in the video based on the sensitive region information contained in the video. Thus, the above-mentioned corresponding processing can be realized based on the above-mentioned rule information in the first smart contract. For details, please refer to the above-mentioned related content, which will not be repeated here.

[0088] In step S510, based on the first smart contract, the desensitized live video data stream is displayed, and the attention information of the user in the process of displaying the desensitized live video data stream is detected.

[0089] In implementation, the first smart contract can be provided with related rule information of displaying the desensitized live video data stream and detecting the attention information of the user in the process of displaying the desensitized live video data stream. Thus, the above-mentioned corresponding processing can be realized based on the above-mentioned rule information in the first smart contract. For details, please refer to the above-mentioned related content, which will not be repeated here.

[0090] In step S512, if the attention information of the user indicates that the attention of the user is not in the preset region of the video corresponding to the live video data stream, based on the first smart contract, the attention information of the user and the desensitized live video data stream, the adjustment strategy of the attention of the user is determined, and the adjustment strategy of the attention of the user is output to the live side corresponding to the live video data stream. The adjustment strategy of the attention of the user is used to instruct the live side to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream.

[0091] In implementation, the first smart contract can be configured with the adjustment strategy of user attention based on the user's attention information and the desensitized live video data stream, and output the relevant rule information of the adjustment strategy of user attention to the live party corresponding to the live video data stream. In this way, based on the above-mentioned rule information in the first smart contract, the above-mentioned corresponding processing can be realized. For details, please refer to the above-mentioned related content, which will not be repeated here.

[0092] The specific processing of steps S504 to S512 can be referred to the related content in the above-mentioned embodiment one to embodiment two, that is, the various processing involved in the above-mentioned embodiment one to embodiment two can be realized through the first smart contract.

[0093] In actual application, the processing mode of detecting the user's attention information in the process of displaying the desensitized live video data stream based on the first smart contract in step S510 can include multiple, the following provides an optional processing mode, which can include the following content: based on the first smart contract, input the desensitized live video data stream into the pre-trained attention detection model to obtain the user's attention information in the process of displaying the desensitized live video data stream.

[0094] In implementation, the first smart contract can be configured with the related rule information of inputting the desensitized live video data stream into the pre-trained attention detection model to obtain the user's attention information in the process of displaying the desensitized live video data stream. In this way, based on the above-mentioned rule information in the first smart contract, the above-mentioned corresponding processing can be realized. For details, please refer to the above-mentioned related content, which will not be repeated here.

[0095] The attention detection model can be trained in the following way: based on the third smart contract pre-deployed in the blockchain system to obtain the second training sample carrying the user attention graph, the second training sample is composed of historical live video data stream; based on the third smart contract and the second training sample, and based on the second loss function of mean square error, the attention detection model is trained to obtain the trained attention detection model, and the attention detection model is a model constructed based on image segmentation UNet algorithm.

[0096] In implementation, the third smart contract can be configured with the related rule information of obtaining the second training sample carrying the user attention graph, and training the attention detection model based on the second training sample and based on the second loss function of mean square error. In this way, based on the above-mentioned rule information in the third smart contract, the above-mentioned corresponding processing can be realized. For details, please refer to the above-mentioned related content, which will not be repeated here.

[0097] In actual application, the processing manner of determining the adjustment strategy of the user's attention based on the first smart contract, the attention information of the user and the desensitized live video data stream in step S512 can include multiple manners. An optional processing manner is provided below, which can include the following content: based on the first smart contract, inputting the attention information of the user and the desensitized live video data stream into a pre-trained attention adjustment model to obtain the adjustment strategy of the user's attention.

[0098] In implementation, the first smart contract can be provided with related rule information of inputting the attention information of the user and the desensitized live video data stream into a pre-trained attention adjustment model to obtain the adjustment strategy of the user's attention. Based on the above rule information in the first smart contract, the above corresponding processing can be implemented. For details, refer to the above related content, which will not be repeated here.

[0099] The attention adjustment model can be trained in the following manner: based on the fourth smart contract pre-deployed in the blockchain system, obtain the third training sample carrying the adjustment strategy of the user's attention and the corresponding attention graph, the third training sample is composed of historical live video data stream; based on the fourth smart contract and the third training sample, and a preset Softmax loss function, train the attention adjustment model to obtain the trained attention adjustment model, and the attention adjustment model is a model constructed based on a residual network ResNet algorithm.

[0100] In implementation, the fourth smart contract can be provided with related rule information of obtaining the third training sample carrying the adjustment strategy of the user's attention and the corresponding attention graph, and training the attention adjustment model based on the third training sample and a preset Softmax loss function. Based on the above rule information in the fourth smart contract, the above corresponding processing can be implemented. For details, refer to the above related content, which will not be repeated here.

[0101] In the embodiments of the present specification, the sensitive area information includes information of a sensitive area in a space environment where the live streaming party is located, information of a preset type of object of the live streaming party or a local area of the object, and information related to a live streaming object corresponding to the live video data stream.

[0102] In actual application, the processing manner of step S508 can include multiple manners. An optional processing manner is provided below, which can include the following content: based on the first smart contract and the pre-trained information recognition model, recognizing the sensitive information in the sensitive area corresponding to the sensitive area information contained in the video to obtain the sensitive information in the sensitive area contained in the video; based on the first smart contract, performing desensitization processing on the sensitive information in the sensitive area contained in the video to obtain the desensitized live video data stream.

[0103] In implementation, the first smart contract can be provided with rules information for identifying sensitive information in the sensitive region corresponding to the sensitive region information contained in the video based on the pre-trained information recognition model, obtaining the sensitive information in the sensitive region contained in the video, and desensitizing the sensitive information in the sensitive region contained in the video. Based on the above rules information in the first smart contract, the corresponding processing can be realized. For details, please refer to the above related content, which will not be repeated here.

[0104] The sensitive region detection model can be trained in the following manner: obtaining a first training sample based on a second smart contract pre-deployed in the blockchain system, the first training sample being composed of historical live video data streams, and the video corresponding to the historical live video data streams containing a sensitive region; training the sensitive region detection model based on the second smart contract, the first training sample, and a first loss function based on mean square error, to obtain a trained sensitive region detection model, the sensitive region detection model being a model constructed based on a neural network algorithm, and the neural network algorithm including a FasterRCNN algorithm.

[0105] In implementation, the second smart contract can be provided with rules information for obtaining a first training sample and training a sensitive region detection model based on the first training sample and a first loss function based on mean square error. Based on the above rules information in the second smart contract, the corresponding processing can be realized. For details, please refer to the above related content, which will not be repeated here.

[0106] In the embodiments of the present specification, the adjustment strategy of attention includes a dialogue adjustment strategy and / or a behavior action adjustment strategy.

[0107] In actual application, each model involved above can be stored in the blockchain system or other storage devices. For a model stored in other storage devices, considering that the model may need to be updated regularly or irregularly, since the blockchain system has the feature of being tamper-proof, if the model is stored in the blockchain system, subsequent operations such as frequent uploading, deletion, and identity verification of the uploader of the model in the blockchain system will increase the processing pressure of the blockchain system. In order to improve the processing efficiency and reduce the processing pressure of the blockchain system, the model can be pre-stored in a specified storage address of the storage device, and the storage address (which can be in the form of index information) is uploaded to the blockchain system. Since the storage address can be fixed and stored in the blockchain system, the tamper-proof nature of the data in the blockchain system is ensured, and the model can also be updated regularly or irregularly in the storage device.

[0108] The embodiment of the present specification provides a live video processing method, which is applied to a blockchain system, acquires a live video data stream to be displayed through a first smart contract, inputs the live video data stream into a pre-trained sensitive area detection model to identify sensitive areas contained in a video corresponding to the live video data stream, obtains sensitive area information contained in the video, then, based on the sensitive area information contained in the video, desensitizes sensitive information in the sensitive areas contained in the video to obtain a desensitized live video data stream, displays the desensitized live video data stream, and detects attention information of a user in the process of displaying the desensitized live video data stream. If the attention information of the user indicates that the attention of the user is not in a preset area of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, a user attention adjustment strategy is determined, and the user attention adjustment strategy is output to a live streaming party corresponding to the live video data stream. The user attention adjustment strategy is used to instruct the live streaming party to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream. In this way, the deep learning model (i.e., the sensitive area detection model) is used to convert a large amount of manual work into an automatic, quasi-real-time standardized process. The sensitive area detection model can automatically locate private information / irrelevant information of the live streaming object, and then process it through the desensitization technology without human intervention. In terms of attention attraction, the user's attention can also be automatically detected, and a more optimal attention adjustment strategy can be automatically predicted and provided to the live streaming party without relying on the experience of the live streaming party, thereby improving the live streaming efficiency.

[0109] Embodiment four

[0110] The above is the live video processing method provided by the embodiment of the present specification. Based on the same idea, the embodiment of the present specification also provides a live video processing device, as shown in Figure 6 .

[0111] The live video processing device comprises a data acquisition module 601, a sensitive area identification module 602, a desensitization module 603, an attention detection module 604, and an attention adjustment module 605, wherein:

[0112] The data acquisition module 601 acquires a live video data stream to be displayed;

[0113] The sensitive area identification module 602 inputs the live video data stream into a pre-trained sensitive area detection model to identify sensitive areas contained in a video corresponding to the live video data stream, and obtains sensitive area information contained in the video;

[0114] The desensitization module 603 performs desensitization processing on sensitive information in a sensitive region included in the video based on sensitive region information included in the video, to obtain a desensitized live video data stream.

[0115] The attention detection module 604 displays the desensitized live video data stream and detects attention information of a user in the process of displaying the desensitized live video data stream.

[0116] The attention adjustment module 605 determines an adjustment strategy for the user's attention based on the user's attention information and the desensitized live video data stream if the user's attention information indicates that the user's attention is not in a preset region of a video corresponding to the live video data stream, and outputs the adjustment strategy for the user's attention to a live streaming party corresponding to the live video data stream. The adjustment strategy for the user's attention is used to instruct the live streaming party to guide the user's attention to be in the preset region of the video corresponding to the live video data stream.

[0117] In the embodiments of the present specification, the attention detection module 604 inputs the desensitized live video data stream into a pre-trained attention detection model to obtain the user's attention information in the process of displaying the desensitized live video data stream.

[0118] In the embodiments of the present specification, the attention adjustment module 605 inputs the user's attention information and the desensitized live video data stream into a pre-trained attention adjustment model to obtain the adjustment strategy for the user's attention.

[0119] In the embodiments of the present specification, the sensitive region information includes information of a sensitive region in a space environment where the live streaming party is located, information of a preset type of object or a local region of the object of the live streaming party, and information related to a live streaming object corresponding to the live video data stream.

[0120] In the embodiments of the present specification, the desensitization module 603 includes:

[0121] The sensitive information determination unit identifies sensitive information in a sensitive region corresponding to the sensitive region information included in the video based on a pre-trained information recognition model, to obtain the sensitive information in the sensitive region included in the video.

[0122] The desensitization unit performs desensitization processing on the sensitive information in the sensitive region included in the video, to obtain a desensitized live video data stream.

[0123] In the embodiments of the present specification, further comprising:

[0124] The first sample acquisition module acquires a first training sample, the first training sample is composed of a historical live video data stream, and the historical live video data stream corresponds to a video containing a sensitive area;

[0125] The first training module trains the sensitive area detection model based on the first training sample and a first loss function based on a mean square error, obtains a trained sensitive area detection model, the sensitive area detection model is a model constructed based on a neural network algorithm, and the neural network algorithm includes a FasterRCNN algorithm.

[0126] In the embodiments of the present specification, the following are further included:

[0127] The second sample acquisition module acquires a second training sample carrying a user attention graph, and the second training sample is composed of a historical live video data stream;

[0128] The second training module trains the attention detection model based on the second training sample and a second loss function based on a mean square error, obtains a trained attention detection model, and the attention detection model is a model constructed based on an image segmentation UNet algorithm.

[0129] In the embodiments of the present specification, the following are further included:

[0130] The third sample acquisition module acquires a third training sample carrying an adjustment strategy of user attention and a corresponding attention graph, and the third training sample is composed of a historical live video data stream;

[0131] The third training module trains the attention adjustment model based on the third training sample and a preset Softmax loss function, obtains a trained attention adjustment model, and the attention adjustment model is a model constructed based on a residual network ResNet algorithm.

[0132] In the embodiments of the present specification, the adjustment strategy of the attention includes a dialogue adjustment strategy and / or a behavior action adjustment strategy.

[0133] The embodiment of the present specification provides a live video processing apparatus, acquires a live video data stream to be displayed, inputs the live video data stream into a pre-trained sensitive area detection model to identify a sensitive area contained in a video corresponding to the live video data stream, obtains sensitive area information contained in the video, then, based on the sensitive area information contained in the video, desensitizes sensitive information in the sensitive area contained in the video to obtain a desensitized live video data stream, displays the desensitized live video data stream, and detects attention information of a user in the process of displaying the desensitized live video data stream, if the attention information of the user indicates that the attention of the user is not in a preset area of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, determines an adjustment strategy of the attention of the user, and outputs the adjustment strategy of the attention of the user to a live party corresponding to the live video data stream, the adjustment strategy of the attention of the user is used to instruct the live party to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream, in this way, the deep learning model (i.e. the sensitive area detection model) is used to convert a large amount of manual work into an automatic, quasi-real-time standardized process, and the sensitive area detection model can automatically locate to private information / information irrelevant to the live object, then the desensitization technology is used for processing, without manual intervention, and in terms of attention attraction, the attention of the user can also be automatically detected, and a more optimal adjustment strategy of the attention is automatically predicted and given to the live party, without relying on the experience of the live party, improving the live efficiency.

[0134] Embodiment five

[0135] Based on the same idea, the embodiment of the present specification also provides a live video processing apparatus, which is a device in a blockchain system, as shown in Figure 7

[0136] The live video processing apparatus comprises a contract deployment module 701, a data acquisition module 702, a sensitive area identification module 703, a desensitization module 704, an attention detection module 705 and an attention adjustment module 706, wherein:

[0137] The contract deployment module 701 acquires rule information for processing a live video, generates a corresponding first smart contract based on the rule information for processing the live video, and deploys the first smart contract into the blockchain system;

[0138] The data acquisition module 702 acquires a live video data stream to be displayed based on the first smart contract;

[0139] ​The sensitive area identification module 703 inputs the live video data stream into a pre-trained sensitive area detection model based on the first smart contract, to identify sensitive areas contained in a video corresponding to the live video data stream, to obtain sensitive area information contained in the video.

[0140] The desensitization module 704 performs desensitization processing on sensitive information in the sensitive areas contained in the video based on the first smart contract and the sensitive area information contained in the video, to obtain a desensitized live video data stream.

[0141] The attention detection module 705 displays the desensitized live video data stream based on the first smart contract, and detects attention information of a user in the process of displaying the desensitized live video data stream.

[0142] The attention adjustment module 706 determines an adjustment strategy for the user's attention based on the first smart contract, the user's attention information, and the desensitized live video data stream, if the user's attention information indicates that the user's attention is not in a preset area of a video corresponding to the live video data stream, and outputs the adjustment strategy for the user's attention to a live streaming party corresponding to the live video data stream. The adjustment strategy for the user's attention is used to instruct the live streaming party to guide the user's attention to be in the preset area of the video corresponding to the live video data stream.

[0143] In the embodiments of the present specification, the apparatus further comprises:

[0144] The first sample acquisition module acquires a first training sample based on a second smart contract pre-deployed in the blockchain system, the first training sample being composed of historical live video data streams, and a video corresponding to the historical live video data streams containing sensitive areas.

[0145] The first training module trains the sensitive area detection model based on the second smart contract, the first training sample, and a first loss function based on mean square error, to obtain a trained sensitive area detection model, the sensitive area detection model being a model constructed based on a neural network algorithm, and the neural network algorithm including a FasterRCNN algorithm.

[0146] The embodiment of the present specification provides a live video processing apparatus, acquires a live video data stream to be displayed, inputs the live video data stream into a pre-trained sensitive area detection model to identify sensitive areas contained in a video corresponding to the live video data stream, obtains sensitive area information contained in the video, then, based on the sensitive area information contained in the video, desensitizes sensitive information in the sensitive areas contained in the video to obtain a desensitized live video data stream, displays the desensitized live video data stream, and detects user attention information in the process of displaying the desensitized live video data stream. If the user attention information indicates that the user's attention is not in a preset area of the video corresponding to the live video data stream, based on the user attention information and the desensitized live video data stream, determine the adjustment strategy of the user's attention, and output the adjustment strategy of the user's attention to the live video data stream corresponding to the live party. The adjustment strategy of the user's attention is used to instruct the live party to guide the user's attention to be in the preset area of the video corresponding to the live video data stream. In this way, the deep learning model (i.e. sensitive area detection model) is used to convert a large amount of manual work into an automatic, quasi-real-time standardized process. The sensitive area detection model can automatically locate private information / irrelevant information of the live object, and then process it through desensitization technology without human intervention. In terms of attention attraction, the user's attention can also be automatically detected, and the optimal adjustment strategy of the attention is automatically predicted and provided to the live party, without relying on the experience of the live party. The live efficiency is improved.

[0147] Embodiment six

[0148] The above is the live video processing apparatus provided by the embodiment of the present specification. Based on the same idea, the embodiment of the present specification also provides a live video processing device. As shown in the Figure 8 The live video processing device can be a server or a device in a blockchain system provided in the above embodiments.

[0149] The live video processing device can have a large difference due to different configurations or performances, and can include one or more processors 801 and memories 802, and the memories 802 can store one or more stored applications or data. Among them, the memories 802 can be temporary storage or persistent storage. The applications stored in the memories 802 can include one or more modules (not shown in the figure), and each module can include a series of computer executable instructions in the live video processing device. Further, the processor 801 can be configured to communicate with the memory 802 and execute a series of computer executable instructions in the memory 802 on the live video processing device. The live video processing device can also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, and one or more keyboards 806.

[0150] In particular, in the embodiment, the live video processing device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and the one or more programs can include one or more modules, and each module can include a series of computer executable instructions in the live video processing device, and the one or more processors are configured to execute the one or more programs include the following computer executable instructions:

[0151] Obtain a live video data stream to be displayed;

[0152] Input the live video data stream into a pre-trained sensitive region detection model to identify sensitive regions contained in a video corresponding to the live video data stream, and obtain sensitive region information contained in the video;

[0153] Based on the sensitive region information contained in the video, desensitize the sensitive information in the sensitive region contained in the video to obtain a desensitized live video data stream;

[0154] Display the desensitized live video data stream, and detect the attention information of the user in the process of displaying the desensitized live video data stream;

[0155] If the user's attention information indicates that the user's attention is not in a preset region of the video corresponding to the live video data stream, based on the user's attention information and the desensitized live video data stream, determine the adjustment strategy of the user's attention, and output the adjustment strategy of the user's attention to the live party corresponding to the live video data stream, the adjustment strategy of the user's attention is used to instruct the live party to guide the user's attention to be in the preset region of the video corresponding to the live video data stream.

[0156] In the embodiments of the present specification, the detection of the attention information of the user in the process of displaying the desensitized live video data stream includes:

[0157] The desensitized live video data stream is input into the pre-trained attention detection model to obtain the attention information of the user in the process of displaying the desensitized live video data stream.

[0158] In the embodiments of the present specification, the determination of the adjustment strategy of the user's attention based on the attention information of the user and the desensitized live video data stream includes:

[0159] The attention information of the user and the desensitized live video data stream are input into the pre-trained attention adjustment model to obtain the adjustment strategy of the user's attention.

[0160] In the embodiments of the present specification, the sensitive region information includes information of a sensitive region in a space environment where the live party is located, information of a preset type of object of the live party or a local region of the object, and information related to a live object corresponding to the live video data stream.

[0161] In the embodiments of the present specification, the desensitization processing of the sensitive information in the sensitive region included in the video based on the sensitive region information included in the video to obtain the desensitized live video data stream includes:

[0162] The sensitive information in the sensitive region corresponding to the sensitive region information included in the video is identified based on the pre-trained information identification model to obtain the sensitive information in the sensitive region included in the video.

[0163] The sensitive information in the sensitive region included in the video is desensitized to obtain the desensitized live video data stream.

[0164] In the embodiments of the present specification, it also includes:

[0165] The first training sample is obtained, the first training sample is composed of historical live video data stream, and the video corresponding to the historical live video data stream includes a sensitive region.

[0166] The sensitive region detection model is trained based on the first training sample and a first loss function based on mean square error to obtain the trained sensitive region detection model, the sensitive region detection model is a model constructed based on a neural network algorithm, and the neural network algorithm includes a FasterRCNN algorithm.

[0167] In the embodiments of the present specification, it also includes:

[0168] obtain a second training sample carrying a user attention map, the second training sample being composed of historical live video data streams;

[0169] train the attention detection model based on the second training sample and a second loss function based on mean square error, to obtain a trained attention detection model, the attention detection model being a model constructed based on an image segmentation UNet algorithm.

[0170] In the embodiments of the present specification, the method further comprises:

[0171] obtain a third training sample carrying an adjustment strategy of user attention and a corresponding attention map, the third training sample being composed of historical live video data streams;

[0172] train the attention adjustment model based on the third training sample and a preset Softmax loss function, to obtain a trained attention adjustment model, the attention adjustment model being a model constructed based on a residual network ResNet algorithm.

[0173] In the embodiments of the present specification, the adjustment strategy of attention includes a dialogue adjustment strategy and / or a behavior action adjustment strategy.

[0174] In particular, in the present embodiment, the live video processing device comprises a memory and one or more programs, wherein one or more programs are stored in the memory, and the one or more programs can include one or more modules, and each module can include a series of computer executable instructions in the live video processing device, and the one or more programs configured to be executed by one or more processors include computer executable instructions for:

[0175] obtain rule information for processing live video, generate a corresponding first smart contract based on the rule information for processing live video, and deploy the first smart contract to the blockchain system;

[0176] obtain a live video data stream to be displayed based on the first smart contract;

[0177] based on the first smart contract, input the live video data stream into a pre-trained sensitive area detection model to identify sensitive areas contained in the video corresponding to the live video data stream, to obtain sensitive area information contained in the video;

[0178] based on the first smart contract and the sensitive area information contained in the video, perform desensitization processing on sensitive information in the sensitive area contained in the video, to obtain a desensitized live video data stream;

[0179] based on the first smart contract, display the desensitized live video data stream, and detect attention information of the user in the process of displaying the desensitized live video data stream;

[0180] If the attention information of the user indicates that the attention of the user is not in the preset area of the video corresponding to the live video data stream, based on the first smart contract, the attention information of the user and the desensitized live video data stream, determine the adjustment strategy of the attention of the user, and output the adjustment strategy of the attention of the user to the live side corresponding to the live video data stream, the adjustment strategy of the attention of the user is used to instruct the live side to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream.

[0181] The embodiments of the present specification also include:

[0182] based on the second smart contract pre-deployed in the blockchain system to obtain a first training sample, the first training sample is composed of historical live video data stream, the video corresponding to the historical live video data stream contains a sensitive area;

[0183] based on the second smart contract, the first training sample, and a first loss function based on mean square error, the sensitive area detection model is trained to obtain a trained sensitive area detection model, the sensitive area detection model is a model constructed based on a neural network algorithm, and the neural network algorithm includes a FasterRCNN algorithm.

[0184] The embodiment of the specification provides a live video processing device, acquires a live video data stream to be displayed, inputs the live video data stream into a pre-trained sensitive area detection model to identify a sensitive area contained in a video corresponding to the live video data stream, obtains sensitive area information contained in the video, then, based on the sensitive area information contained in the video, desensitizes sensitive information in the sensitive area contained in the video to obtain a desensitized live video data stream, displays the desensitized live video data stream, and detects attention information of a user in the process of displaying the desensitized live video data stream, if the attention information of the user indicates that the attention of the user is not in a preset area of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, determines an adjustment strategy of the attention of the user, and outputs the adjustment strategy of the attention of the user to a live side corresponding to the live video data stream, the adjustment strategy of the attention of the user is used to instruct the live side to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream, in this way, the deep learning model (i.e. the sensitive area detection model) is used to convert a large amount of manual work into an automatic, quasi-real-time standardized process, and the sensitive area detection model can automatically locate to private information / information irrelevant to the live object, and then the desensitization technology is used for processing without manual intervention, and in the aspect of attention attraction, the attention of the user can also be automatically detected, and a more optimal adjustment strategy of the attention is automatically predicted and given to the live side without relying on the experience of the live side, thereby improving the live efficiency.

[0185] Embodiment seven

[0186] Further, based on the above Figures 1 to 5B The one or more embodiments of the specification also provide a storage medium for storing computer executable instruction information, in a specific embodiment, the storage medium can be a U disk, an optical disk, a hard disk, etc., and the computer executable instruction information stored in the storage medium can realize the following processes when executed by a processor:

[0187] Acquire a live video data stream to be displayed;

[0188] Input the live video data stream into a pre-trained sensitive area detection model to identify a sensitive area contained in a video corresponding to the live video data stream, and obtain sensitive area information contained in the video;

[0189] Based on the sensitive area information contained in the video, desensitize sensitive information in the sensitive area contained in the video to obtain a desensitized live video data stream;

[0190] display the desensitized live video data stream, and detect attention information of the user in the process of displaying the desensitized live video data stream;

[0191] If the attention information of the user indicates that the attention of the user is not in the preset region of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, a user attention adjustment strategy is determined, and the user attention adjustment strategy is output to the live side corresponding to the live video data stream, the user attention adjustment strategy is used to instruct the live side to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream.

[0192] In the embodiments of the present specification, the detection of the attention information of the user in the process of displaying the desensitized live video data stream comprises:

[0193] The desensitized live video data stream is input into a pre-trained attention detection model to obtain the attention information of the user in the process of displaying the desensitized live video data stream.

[0194] In the embodiments of the present specification, the determination of the user attention adjustment strategy based on the attention information of the user and the desensitized live video data stream comprises:

[0195] The user's attention information and the desensitized live video data stream are input into a pre-trained attention adjustment model to obtain the user's attention adjustment strategy.

[0196] In the embodiments of the present specification, the sensitive region information includes information of a sensitive region in a space environment where the live side is located, information of a preset type of object or a local region of the object of the live side, and information related to a live object corresponding to the live video data stream.

[0197] In the embodiments of the present specification, the desensitization processing of the sensitive information in the sensitive region included in the video based on the sensitive region information included in the video to obtain the desensitized live video data stream comprises:

[0198] Based on a pre-trained information recognition model, the sensitive information in the sensitive region corresponding to the sensitive region information included in the video is recognized to obtain the sensitive information in the sensitive region included in the video.

[0199] The sensitive information in the sensitive region included in the video is desensitized to obtain the desensitized live video data stream.

[0200] In the embodiments of the present specification, it further comprises:

[0201] obtain a first training sample, the first training sample being composed of a historical live video data stream, the historical live video data stream corresponding to a video containing a sensitive area;

[0202] train the sensitive area detection model based on the first training sample and a first loss function based on mean square error, to obtain a trained sensitive area detection model, the sensitive area detection model being a model constructed based on a neural network algorithm, the neural network algorithm including a FasterRCNN algorithm.

[0203] In the embodiments of the present specification, the following are further included:

[0204] obtain a second training sample carrying a user attention map, the second training sample being composed of a historical live video data stream;

[0205] train the attention detection model based on the second training sample and a second loss function based on mean square error, to obtain a trained attention detection model, the attention detection model being a model constructed based on an image segmentation UNet algorithm.

[0206] In the embodiments of the present specification, the following are further included:

[0207] obtain a third training sample carrying an adjustment strategy of user attention and a corresponding attention map, the third training sample being composed of a historical live video data stream;

[0208] train the attention adjustment model based on the third training sample and a preset Softmax loss function, to obtain a trained attention adjustment model, the attention adjustment model being a model constructed based on a residual network ResNet algorithm.

[0209] In the embodiments of the present specification, the adjustment strategy of the attention includes a dialogue adjustment strategy and / or a behavior action adjustment strategy.

[0210] In addition, in another specific embodiment, the storage medium can be a U disk, an optical disk, a hard disk, etc., and the computer executable instruction information stored in the storage medium, when executed by the processor, can implement the following flow:

[0211] obtain rule information for processing a live video, generate a corresponding first smart contract based on the rule information for processing the live video, and deploy the first smart contract to the blockchain system;

[0212] obtain a live video data stream to be displayed based on the first smart contract;

[0213] input the live video data stream into a pre-trained sensitive region detection model based on the first smart contract, to identify sensitive regions contained in a video corresponding to the live video data stream, to obtain sensitive region information contained in the video;

[0214] perform desensitization processing on sensitive information in the sensitive regions contained in the video based on the first smart contract and the sensitive region information contained in the video, to obtain a desensitized live video data stream;

[0215] based on the first smart contract, display the desensitized live video data stream, and detect attention information of a user in a process of displaying the desensitized live video data stream;

[0216] if the attention information of the user indicates that the attention of the user is not in a preset region of the video corresponding to the live video data stream, determine an adjustment strategy for the attention of the user based on the first smart contract, the attention information of the user, and the desensitized live video data stream, and output the adjustment strategy for the attention of the user to a live streaming party corresponding to the live video data stream, the adjustment strategy for the attention of the user being used to instruct the live streaming party to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream.

[0217] In the embodiments of the present specification, the following are further included:

[0218] obtain a first training sample based on a second smart contract pre-deployed in the blockchain system, the first training sample being composed of historical live video data streams, and the video corresponding to the historical live video data streams containing sensitive regions;

[0219] train the sensitive region detection model based on the second smart contract, the first training sample, and a first loss function based on mean square error, to obtain a trained sensitive region detection model, the sensitive region detection model being a model constructed based on a neural network algorithm, and the neural network algorithm including a FasterRCNN algorithm.

[0220] The embodiment of the specification provides a storage medium, acquires a live video data stream to be displayed, inputs the live video data stream into a pre-trained sensitive area detection model to identify a sensitive area contained in a video corresponding to the live video data stream, obtains sensitive area information contained in the video, then, based on the sensitive area information contained in the video, performs desensitization processing on sensitive information in the sensitive area contained in the video, obtains a desensitized live video data stream, displays the desensitized live video data stream, and detects attention information of a user in the process of displaying the desensitized live video data stream, if the attention information of the user indicates that the attention of the user is not in a preset area of the video corresponding to the live video data stream, based on the attention information of the user and the desensitized live video data stream, determines an adjustment strategy of the attention of the user, and outputs the adjustment strategy of the attention of the user to a live side corresponding to the live video data stream, the adjustment strategy of the attention of the user is used to instruct the live side to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream, in this way, the deep learning model (i.e., the sensitive area detection model) is used to convert a large amount of manual work into an automatic, quasi-real-time standardized process, and the sensitive area detection model can automatically locate to private information / irrelevant information of the live object, and then the desensitization technology is used for processing without manual intervention, and in terms of attention attraction, the attention of the user can also be automatically detected, and a more optimal adjustment strategy of the attention is automatically predicted and given to the live side, without relying on the experience of the live side to improve the live efficiency.

[0221] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in the embodiments and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.

[0222] In the 1990s, it was possible to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has advanced, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain a corresponding hardware circuit structure by programming the improved method flow into a hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it, rather than by asking a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating an integrated circuit chip, this programming is now mostly implemented using "logic compiler" software, which is similar to software compilers used in program development, and the original code to be compiled is written in a specific programming language, which is called a hardware description language (HDL), and there are many such languages, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should be aware that, as long as the method flow is logically programmed in one of the above hardware description languages and programmed into an integrated circuit, a hardware circuit implementing the logical method flow can be easily obtained.

[0223] The controller can be implemented in any suitable way, e.g. the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, e.g. software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of controllers include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91 SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to being implemented in pure computer readable program code form, the controller can perfectly well be implemented by means of logic programmed into logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers, etc. to perform the same functions. The controller can thus be considered as a hardware component, and the means comprised therein for performing various functions can be considered as structures within the hardware component. Alternatively, or even, the means for performing various functions can be considered as both a software module implementing a method and a structure within a hardware component.

[0224] The systems, apparatuses, modules or units illustrated by the above embodiments can be implemented by computer chips or entities, or products with certain functions. A typical implementation device is a computer. Specifically, the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0225] For the sake of description, the above apparatuses are described in various units with functions respectively. Of course, the functions of each unit can be implemented in one or more software and / or hardware in implementing one or more embodiments of the present specification.

[0226] Those skilled in the art will understand that the embodiments of the present specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0227] The embodiments of the present specification are described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable electronic devices to produce a machine, so that the instructions executed by the computer or other programmable electronic devices generate a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus.

[0228] These computer program instructions can also be stored in a computer readable memory that can direct the computer or other programmable electronic devices to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus.

[0229] These computer program instructions can also be loaded into the computer or other programmable electronic devices, so that a series of operation steps are performed on the computer or other programmable electronic devices to produce a computer implemented process, so that the instructions executed on the computer or other programmable electronic devices provide steps for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus. Figure 1 The functions of one or more flows and / or blocks in the flowcharts and / or block diagrams can be implemented by an apparatus.

[0230] In a typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces and memories.

[0231] The memory can include non-persistent memory in the computer readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of the computer readable medium.

[0232] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0233] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, such that processes, methods, articles or devices that comprise a list of elements not only include those elements, but also include other elements not expressly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or device that includes the element.

[0234] Those skilled in the art will appreciate that embodiments of the present specification can be provided as methods, systems or computer program products. Therefore, one or more embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0235] One or more embodiments of the present specification can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. One or more embodiments of the present specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are connected through a communication network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including storage devices.

[0236] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the system embodiments, the description is relatively simple because the system embodiments are basically similar to the method embodiments, and the relevant parts can be referred to the part of the method embodiments.

[0237] The above only describes the embodiments of the specification and is not intended to limit the specification. The specification can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the specification shall be included in the scope of claims of the specification.

Claims

1. A method for processing a live video, the method comprising: obtaining a live video data stream to be displayed; inputting the live video data stream into a pre-trained sensitive area detection model to identify sensitive areas contained in a video corresponding to the live video data stream, to obtain sensitive area information contained in the video, the sensitive areas including a sensitive area in a spatial environment where a live streaming party is located, a local area of a preset type of object of the live streaming party or the object, and an area where related information of a live streaming object corresponding to the live video data stream is located; performing desensitization processing on sensitive information in the sensitive areas contained in the video based on the sensitive area information contained in the video, to obtain a desensitized live video data stream; displaying the desensitized live video data stream, and inputting the desensitized live video data stream into a pre-trained attention detection model to obtain attention information of a user in a process of displaying the desensitized live video data stream, the attention detection model being trained based on a second training sample constituted by historical live video data streams carrying a user attention map; if the attention information of the user indicates that the attention of the user is not in a preset area of the video corresponding to the live video data stream, inputting the attention information of the user and the desensitized live video data stream into a pre-trained attention adjustment model to obtain an adjustment strategy of the attention of the user, and outputting the adjustment strategy of the attention of the user to the live streaming party corresponding to the live video data stream, the attention adjustment model being trained based on a third training sample constituted by historical live video data streams carrying an adjustment strategy of the attention of the user and an attention map corresponding to the adjustment strategy of the attention of the user, the adjustment strategy of the attention of the user being used to instruct the live streaming party to guide the attention of the user to be in the preset area of the video corresponding to the live video data stream, and the attention of the user being adjusted by text provided in the adjustment strategy of the attention of the user or by a combination of multiple types of images, videos, audios, and text for providing the text in the adjustment strategy of the attention of the user.

2. The method of claim 1, wherein the sensitive area information includes information of a sensitive area in a spatial environment where the live streaming party is located, information of a local area of a preset type of object of the live streaming party or the object, and information related to the live streaming object corresponding to the live video data stream.

3. The method of claim 1, wherein the desensitization processing on the sensitive information in the sensitive areas contained in the video based on the sensitive area information contained in the video comprises: identifying the sensitive information in the sensitive areas corresponding to the sensitive area information contained in the video based on a pre-trained information identification model, to obtain the sensitive information in the sensitive areas contained in the video; and performing desensitization processing on the sensitive information in the sensitive areas contained in the video, to obtain the desensitized live video data stream. ​ ​ ​ ​ ​ ​ ​ 4. The method of claim 1, further comprising: obtaining a first training sample constituted by a historical live video data stream, the historical live video data stream corresponding to a video containing a sensitive region; training the sensitive region detection model based on the first training sample and a first loss function based on mean square error, to obtain a trained sensitive region detection model, the sensitive region detection model being a model constructed based on a neural network algorithm, the neural network algorithm including a Faster RCNN algorithm.

5. The method of claim 1, further comprising: obtaining a second training sample carrying a user attention map, the second training sample being constituted by a historical live video data stream; training the attention detection model based on the second training sample and a second loss function based on mean square error, to obtain a trained attention detection model, the attention detection model being a model constructed based on an image segmentation UNet algorithm.

6. The method of claim 1, further comprising: obtaining a third training sample carrying an adjustment strategy of user attention and a corresponding attention map, the third training sample being constituted by a historical live video data stream; training the attention adjustment model based on the third training sample and a preset Softmax loss function, to obtain a trained attention adjustment model, the attention adjustment model being a model constructed based on a residual network ResNet algorithm.

7. The method of claim 1, wherein the adjustment strategy of user attention includes a dialogue adjustment strategy and / or a behavior action adjustment strategy.

8. A live video processing method applied to a blockchain system, the method comprising: obtaining rule information for processing a live video, generating a corresponding first smart contract based on the rule information for processing the live video, and deploying the first smart contract to the blockchain system; obtaining a live video data stream to be displayed based on the first smart contract; inputting the live video data stream into a pre-trained sensitive region detection model based on the first smart contract, to identify a sensitive region contained in a video corresponding to the live video data stream, to obtain sensitive region information contained in the video, the sensitive region including a sensitive region in a space environment where a live party is located, a preset type of an object of the live party or a local region of the object, and a region where related information of a live object corresponding to the live video data stream is located; performing desensitization processing on sensitive information in the sensitive region contained in the video based on the first smart contract and the sensitive region information contained in the video, to obtain a desensitized live video data stream. based on the first smart contract, the desensitized live video data stream is displayed, and the desensitized live video data stream is input into a pre-trained attention detection model to obtain attention information of a user in the process of displaying the desensitized live video data stream, the attention detection model is obtained based on a second training sample composed of historical live video data streams carrying a user attention graph; If the user's attention information indicates that the user's attention is not in the preset area of the video corresponding to the live video data stream, the user's attention information and the desensitized live video data stream are input into a pre-trained attention adjustment model to obtain an adjustment strategy for the user's attention, and the adjustment strategy for the user's attention is output to the live party corresponding to the live video data stream, the attention adjustment model is obtained based on a third training sample composed of historical live video data streams carrying an adjustment strategy for the user's attention and its corresponding attention graph, the adjustment strategy for the user's attention is used to instruct the live party to guide the user's attention to be in the preset area of the video corresponding to the live video data stream, and the user's attention is adjusted through the text provided in the adjustment strategy for the user's attention. The content of the rhetoric or the user's attention is adjusted through the combination of multiple forms of images, videos, audios and texts for providing rhetoric content in the adjustment strategy for the user's attention.

9. The method of claim 8, further comprising: based on a second smart contract pre-deployed in the blockchain system, a first training sample is obtained, the first training sample is composed of historical live video data streams, and the historical live video data streams correspond to videos containing sensitive areas; based on the second smart contract, the first training sample, and a first loss function based on mean square error, the sensitive area detection model is trained to obtain a trained sensitive area detection model, the sensitive area detection model is a model constructed based on a neural network algorithm, and the neural network algorithm includes a Faster RCNN algorithm.

10. A live video processing apparatus, the apparatus comprising: a data acquisition module for acquiring a live video data stream to be displayed; a sensitive area identification module for inputting the live video data stream into a pre-trained sensitive area detection model to identify sensitive areas contained in a video corresponding to the live video data stream, and obtaining sensitive area information contained in the video, the sensitive areas including sensitive areas in a space environment where a live party is located, a preset type of object or a local area of the object of the live party, and an area where related information of a live object corresponding to the live video data stream is located; a desensitization module for desensitizing sensitive information in the sensitive areas contained in the video based on the sensitive area information contained in the video, and obtaining a desensitized live video data stream. The attention detection module displays the desensitized live video data stream and inputs the desensitized live video data stream into a pre-trained attention detection model to obtain attention information of a user in a process of displaying the desensitized live video data stream. The attention detection model is obtained by training based on a second training sample composed of historical live video data streams carrying a user attention graph; The attention adjustment module inputs the attention information of the user and the desensitized live video data stream into a pre-trained attention adjustment model if the attention information of the user indicates that the attention of the user is not in a preset region of a video corresponding to the live video data stream, obtains an adjustment strategy of the attention of the user, and outputs the adjustment strategy of the attention of the user to a live party corresponding to the live video data stream. The attention adjustment model is obtained by training based on a third training sample composed of historical live video data streams carrying an adjustment strategy of the attention of the user and an attention graph corresponding to the adjustment strategy of the attention of the user. The adjustment strategy of the attention of the user is used to instruct the live party to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream. The attention of the user is adjusted by text provided in the adjustment strategy of the attention of the user or by a combination of multiple types of images, videos, audio, and text for providing the text content in the adjustment strategy of the attention of the user.

11. A live video processing apparatus applied to a blockchain system, the apparatus comprising: A contract deployment module acquires rule information for processing a live video, generates a corresponding first smart contract based on the rule information for processing the live video, and deploys the first smart contract into the blockchain system; A data acquisition module acquires a live video data stream to be displayed based on the first smart contract; A sensitive region identification module inputs the live video data stream into a pre-trained sensitive region detection model based on the first smart contract to identify a sensitive region contained in a video corresponding to the live video data stream, and obtains sensitive region information contained in the video. The sensitive region includes a sensitive region in a space environment where a live party is located, a preset type of object or a local region of the object of the live party, and a region where related information of a live object corresponding to the live video data stream is located; A desensitization module desensitizes sensitive information in a sensitive region contained in the video based on the first smart contract and the sensitive region information contained in the video to obtain a desensitized live video data stream. an attention detection module, based on the first smart contract, displaying the desensitized live video data stream and inputting the desensitized live video data stream into a pre-trained attention detection model to obtain attention information of a user in a process of displaying the desensitized live video data stream, the attention detection model being obtained based on a second training sample composed of historical live video data streams carrying user attention maps; an attention adjustment module, if the attention information of the user indicates that the attention of the user is not in a preset region of a video corresponding to the live video data stream, based on the first smart contract, inputting the attention information of the user and the desensitized live video data stream into a pre-trained attention adjustment model to obtain an adjustment strategy of the attention of the user, and outputting the adjustment strategy of the attention of the user to a live streaming party corresponding to the live video data stream, the attention adjustment model being obtained based on a third training sample composed of historical live video data streams carrying adjustment strategies of user attention and corresponding attention maps, the adjustment strategy of the attention of the user being used to instruct the live streaming party to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream, and the attention of the user being adjusted by text provided in the adjustment strategy of the attention of the user or by a combination of multiple types of images, videos, audios and texts for providing text content in the adjustment strategy of the attention of the user.

12. A live video processing device, the live video processing device comprising: a processor; and a memory arranged to store computer-executable instructions that, when executed, cause the processor to: obtain a live video data stream to be displayed; input the live video data stream into a pre-trained sensitive region detection model to identify a sensitive region contained in a video corresponding to the live video data stream, to obtain sensitive region information contained in the video, the sensitive region including a sensitive region in a spatial environment where a live streaming party is located, a preset type of object or a local region of the object of the live streaming party, and a region where related information of a live streaming object corresponding to the live video data stream is located; based on the sensitive region information contained in the video, desensitize sensitive information in the sensitive region contained in the video to obtain a desensitized live video data stream; display the desensitized live video data stream and input the desensitized live video data stream into a pre-trained attention detection model to obtain attention information of a user in a process of displaying the desensitized live video data stream, the attention detection model being obtained based on a second training sample composed of historical live video data streams carrying user attention maps; If the attention information of the user indicates that the attention of the user is not in the preset region of the video corresponding to the live video data stream, the attention information of the user and the desensitized live video data stream are input into a pre-trained attention adjustment model to obtain an adjustment strategy of the user's attention, and the adjustment strategy of the user's attention is output to a live party corresponding to the live video data stream. The attention adjustment model is obtained by training based on a third training sample composed of historical live video data streams carrying an adjustment strategy of user attention and a corresponding attention atlas, the adjustment strategy of the user's attention is used to instruct the live party to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream, and the attention of the user is adjusted by text provided in the adjustment strategy of the user's attention or by a combination of multiple types of images, videos, audio and text for providing the content of the text in the adjustment strategy of the user's attention.

13. A live video processing device, the device being a device in a blockchain system, the live video processing device comprising: a processor; and a memory arranged to store computer executable instructions that, when executed, cause the processor to: obtain rule information for processing a live video, generate a corresponding first smart contract based on the rule information for processing the live video, and deploy the first smart contract to the blockchain system; obtain a live video data stream to be displayed based on the first smart contract; input the live video data stream into a pre-trained sensitive region detection model based on the first smart contract to identify sensitive regions contained in a video corresponding to the live video data stream, and obtain sensitive region information contained in the video, the sensitive regions including sensitive regions in a space environment where a live party is located, a preset type of object or a local region of the object of the live party, and a region where related information of a live object corresponding to the live video data stream is located; perform desensitization processing on sensitive information in the sensitive regions contained in the video based on the first smart contract and the sensitive region information contained in the video, and obtain a desensitized live video data stream; based on the first smart contract, display the desensitized live video data stream, and input the desensitized live video data stream into a pre-trained attention detection model to obtain attention information of a user in a process of displaying the desensitized live video data stream, the attention detection model being obtained by training based on a second training sample composed of historical live video data streams carrying a user attention atlas; If the attention information of the user indicates that the attention of the user is not in the preset region of the video corresponding to the live video data stream, based on the first smart contract, the attention information of the user and the desensitized live video data stream are input into a pre-trained attention adjustment model to obtain an adjustment strategy of the user's attention, and the adjustment strategy of the user's attention is output to a live party corresponding to the live video data stream. The attention adjustment model is obtained by training based on a third training sample composed of historical live video data streams carrying an adjustment strategy of user attention and a corresponding attention graph, and the adjustment strategy of the user's attention is used to instruct the live party to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream. The attention of the user is adjusted by the text provided in the rhetoric content in the adjustment strategy of the user's attention or by the combination of multiple types of images, videos, audio and text for providing rhetoric content in the adjustment strategy of the user's attention.

14. A storage medium for storing computer executable instructions, the executable instructions, when executed by a processor, implement the following flow: obtaining a live video data stream to be displayed; inputting the live video data stream into a pre-trained sensitive region detection model to identify sensitive regions contained in a video corresponding to the live video data stream, to obtain sensitive region information contained in the video, the sensitive regions including sensitive regions in a space environment where a live party is located, a preset type of object or a local region of the object of the live party, and a region where related information of a live object corresponding to the live video data stream is located; based on the sensitive region information contained in the video, desensitizing sensitive information in the sensitive regions contained in the video to obtain a desensitized live video data stream; displaying the desensitized live video data stream and inputting the desensitized live video data stream into a pre-trained attention detection model to obtain attention information of a user in a process of displaying the desensitized live video data stream, the attention detection model being obtained by training based on a second training sample composed of historical live video data streams carrying a user attention graph; If the attention information of the user indicates that the attention of the user is not in the preset region of the video corresponding to the live video stream, the attention information of the user and the desensitized live video stream are input into a pre-trained attention adjustment model to obtain an adjustment strategy of the user's attention, and the adjustment strategy of the user's attention is output to a live party corresponding to the live video stream. The attention adjustment model is obtained by training based on a third training sample composed of historical live video streams carrying an adjustment strategy of user attention and a corresponding attention graph, and the adjustment strategy of the user's attention is used to instruct the live party to guide the attention of the user to be in the preset region of the video corresponding to the live video stream. The attention of the user is adjusted by the text provided in the adjustment strategy of the user's attention, or the attention of the user is adjusted by the combination of multiple types of images, videos, audio and text for providing the content of the text in the adjustment strategy of the user's attention.

15. A storage medium for storing computer executable instructions, the executable instructions, when executed by a processor, implement the following flow: obtain rule information for processing a live video, generate a corresponding first smart contract based on the rule information for processing the live video, and deploy the first smart contract to a blockchain system; based on the first smart contract, obtain a live video data stream to be displayed; based on the first smart contract, input the live video data stream into a pre-trained sensitive region detection model to identify sensitive regions contained in a video corresponding to the live video data stream, and obtain sensitive region information contained in the video, the sensitive region including a sensitive region in a space environment where a live party is located, a predetermined type of object or a local region of the object of the live party, and a region where related information of a live object corresponding to the live video data stream is located; based on the first smart contract and the sensitive region information contained in the video, desensitize sensitive information in the sensitive region contained in the video to obtain a desensitized live video data stream; based on the first smart contract, display the desensitized live video data stream, and input the desensitized live video data stream into a pre-trained attention detection model to obtain attention information of a user in the process of displaying the desensitized live video data stream, the attention detection model being obtained by training based on a second training sample composed of historical live video streams carrying a user attention graph; If the attention information of the user indicates that the attention of the user is not in the preset region of the video corresponding to the live video data stream, based on the first smart contract, the attention information of the user and the desensitized live video data stream are input into a pre-trained attention adjustment model to obtain an adjustment strategy of the user's attention, and the adjustment strategy of the user's attention is output to a live streaming party corresponding to the live video data stream. The attention adjustment model is obtained by training based on a third training sample composed of historical live video data streams carrying an adjustment strategy of user attention and a corresponding attention graph. The adjustment strategy of the user's attention is used to instruct the live streaming party to guide the attention of the user to be in the preset region of the video corresponding to the live video data stream. The attention of the user is adjusted through the text provided in the tactful content in the adjustment strategy of the user's attention or through the combination of multiple kinds of images, videos, audios and texts for providing tactful content in the adjustment strategy of the user's attention.

Citation Information

Patent Citations

  • Data processing method and device, electronic device and storage medium

    CN109558748A

  • Live video processing method and device

    CN111107385A

  • Information interaction method and device

    CN111601064A