Video desensitization methods, systems, devices, and computer storage media
Patent Information
- Application Number
- CN202310306905.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-03-27
AI Technical Summary
这种视频脱敏方式存在很大的缺陷,存在无法准确对视频中的全部敏感信息(例如私人物品及各种账号密码等)进行脱敏的问题
[0015]本申请的技术方案是在至少包括车辆终端和云服务器中,通过车辆终端获取摄像头图像;基于预设的深度学习骨干网络对所述摄像头图像进行编码,得到图像特征,其中,所述图像特征包括车内图像特征和车外图像特征;上传所述图像特征至云服务器,再通过云服务器接收车辆终端上传的图像特征,其中,所述图像特征包括车内图像特征和车外图像特征;基于预设的检测模型对所述图像特征进行检测,得到检测信息,并根据所述检测信息进行场景构建得到脱敏重建场景。通过检测模型对所述图像特征进行检测,得到检测信息,并根据所述检测信息进行场景构建得到脱敏重建场景,进而可以避免无法准确对视频中的全部敏感信息(例如私人物品及各种账号密码等)进行脱敏的现象,本申请的视频脱敏方法可以通过检测模型对所述图像特征进行检测,得到检测信息,并根据所述检测信息进行场景构建得到脱敏重建场景,进而通过逆向思维采集图像中的必要信息进行构建场景,提高了视频脱敏的脱敏准确率。
Smart Images

Figure CN116361854B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, and in particular to a video desensitization method, system, device, and computer storage medium. Background Technology
[0002] With the rapid development of video processing technology, users have increasingly higher requirements for video desensitization. They hope to satisfy the need for desensitization of normal human faces while ensuring the accuracy of desensitization of other sensitive information, which places higher demands on video desensitization technology.
[0003] Traditional video anonymization methods directly anonymize fixed types of sensitive content (such as faces) within the video. This method has a significant drawback: it cannot accurately anonymize all sensitive information in the video (such as personal belongings and various account passwords). In other words, this method suffers from low accuracy due to its inability to comprehensively anonymize all sensitive information in the video. Summary of the Invention
[0004] The main objective of this application is to provide a video desensitization method, system, device, and storage medium, aiming to address the technical problem of improving the desensitization accuracy of video.
[0005] To achieve the above objectives, this application provides a video de-identification method, which is applied to a cloud server. The steps of the video de-identification method include: Receive image features uploaded by the vehicle terminal, wherein the image features include in-vehicle image features and out-of-vehicle image features; The image features are detected based on a preset detection model to obtain detection information, and a desensitized reconstruction scene is obtained by constructing a scene based on the detection information.
[0006] Optionally, the preset detection model includes a 3D human keypoint detection model and a 3D object detection model. The step of detecting the image features based on the preset detection model to obtain detection information includes: Human detection is performed on the image features based on a 3D human key point detection model to obtain human detection information. Object detection is performed on the image features based on a 3D object detection model to obtain object detection information, and detection information is determined based on the human body detection information and the object detection information.
[0007] Optionally, the detection model includes a clothing classification model and an age classification model. The step of performing human detection on the image features based on the 3D human keypoint detection model to obtain human detection information includes: Human image features are determined by performing human detection based on the image features using a 3D human key point detection model. The image features are analyzed using a clothing classification model to obtain first human body feature information, and the image features are analyzed using an age classification model to obtain second human body feature information. Human detection information is obtained by constructing human image features based on the first human feature information, the second human feature information, and preset feature information. The preset feature information includes the initial color features preset by the clothing classification model and the initial age group features preset by the age classification model.
[0008] Optionally, the detection model further includes an animal 3D keypoint model, and the step of determining the detection information based on the human detection information and the object detection information includes: Animal detection is performed on the image features based on the animal 3D key point model to obtain animal detection information, and the human detection information, the object detection information and the animal detection information are summarized as detection information.
[0009] Optionally, the step of constructing a desensitized and reconstructed scene based on the detection information includes: A desensitized and reconstructed scene is obtained by constructing a scene based on at least one of the object detection information, human body detection information, and animal detection information in the detection information and a preset calibration coordinate.
[0010] Optionally, the preset detection model includes an in-vehicle detection model and an out-of-vehicle detection model, and the detection information includes first detection information and second detection information. The step of detecting the image features based on the preset detection model to obtain detection information, and constructing a scene based on the detection information to obtain a desensitized and reconstructed scene includes: The in-vehicle image features are detected based on the in-vehicle detection model to obtain first detection information, and a first desensitized reconstruction scene is obtained by constructing a scene based on the first detection information. The vehicle exterior image features are detected based on the vehicle exterior detection model to obtain second detection information. A second desensitized reconstruction scene is then constructed based on the second detection information. The first desensitized reconstruction scene and the second desensitized reconstruction scene are then combined to form the desensitized reconstruction scene.
[0011] Furthermore, the present invention also provides a video desensitization system, wherein the video desensitization method is applied to a vehicle terminal, and the steps of the video desensitization method include: Acquire camera images; The camera images are encoded based on a pre-set deep learning backbone network to obtain image features, wherein the image features include in-vehicle image features and out-of-vehicle image features. The image features are uploaded to a cloud server, whereby the cloud server determines detection information based on the image features and constructs a de-identified reconstruction scene based on the detection information.
[0012] Furthermore, to achieve the above objectives, the present invention also provides a video desensitization system, the video desensitization system comprising: The vehicle terminal is used to acquire camera images; the camera images are encoded based on a preset deep learning backbone network to obtain image features, wherein the image features include in-vehicle image features and out-of-vehicle image features; the image features are uploaded to a cloud server, wherein the cloud server determines detection information based on the image features and constructs a de-identified reconstruction scene based on the detection information; A cloud server is used to receive image features uploaded by vehicle terminals, wherein the image features include in-vehicle image features and out-of-vehicle image features; the image features are detected based on a preset detection model to obtain detection information, and a desensitized reconstruction scene is obtained by constructing a scene based on the detection information.
[0013] This application also provides a video desensitization device, which includes: a memory, a processor, and a program of the video desensitization method stored in the memory and executable on the processor. When the program of the video desensitization method is executed by the processor, it can implement the steps of the video desensitization method as described above.
[0014] This application also provides a computer storage medium storing a program for implementing a video desensitization method, the program for implementing the video desensitization method being executed by a processor to implement the steps of the video desensitization method as described above.
[0015] The technical solution of this application involves acquiring camera images through a vehicle terminal, which includes at least a vehicle terminal and a cloud server. The camera images are then encoded using a pre-defined deep learning backbone network to obtain image features, including in-vehicle and out-of-vehicle image features. These image features are uploaded to a cloud server, which then receives the uploaded image features from the vehicle terminal. The image features, including in-vehicle and out-of-vehicle image features, are detected using a pre-defined detection model to obtain detection information. Based on this detection information, a scene is constructed to obtain a desensitized and reconstructed scene. By using a detection model to detect the image features and obtain detection information, and then constructing a scene based on this information, the method avoids the inability to accurately desensitize all sensitive information in the video (such as personal items and various account passwords). This video desensitization method uses a detection model to detect the image features, obtain detection information, and construct a scene based on this information. Furthermore, by using reverse thinking to collect necessary information from the image to construct the scene, the accuracy of video desensitization is improved. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the video desensitization device structure in the hardware operating environment involved in the embodiments of the present invention; Figure 2 This is a flowchart illustrating the first embodiment of the video desensitization method of this application; Figure 3 This is a flowchart illustrating the second embodiment of the video desensitization method of this application; Figure 4 This is a schematic diagram of the video desensitization system module of this application; Figure 5 This is a schematic diagram of the video desensitization system structure of this application.
[0019] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0021] Reference Figure 1 , Figure 1 This is a schematic diagram of the video desensitization device structure in the hardware operating environment involved in the embodiments of the present invention.
[0022] like Figure 1 As shown, the video desensitization device may include: a processor 0003, such as a central processing unit (CPU), a communication bus 0001, an acquisition interface 0002, a processing interface 0004, and a memory 0005. The communication bus 0001 is used to establish communication between these components. The acquisition interface 0002 may include an information acquisition system or an acquisition unit such as a computer; optionally, the acquisition interface 0002 may also include a standard wired interface or a wireless interface. The processing interface 0004 may optionally include a standard wired interface or a wireless interface. The memory 0005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. Optionally, the memory 0005 may also be a storage system independent of the aforementioned processor 0003.
[0023] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the video desensitization device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0024] like Figure 1 As shown, the memory 0005, which serves as a storage medium, may include an operating system, an acquisition interface module, an execution interface module, and a video desensitization program.
[0025] exist Figure 1 In the video desensitization device shown, the communication bus 0001 is mainly used to realize the connection and communication between components; the acquisition interface 0002 is mainly used to connect to the backend server and communicate data with the backend server; the processing interface 0004 is mainly used to connect to the deployment end (user end) and communicate data with the deployment end; the processor 0003 and the memory 0005 in the video desensitization device of the present invention can be set in the video desensitization device. The video desensitization device calls the video desensitization program stored in the memory 0005 through the processor 0003 and executes the video desensitization method provided in the embodiment of the present invention.
[0026] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the implementation of a video desensitization method is given first: Traditional video anonymization methods can only anonymize fixed locations or information within an image, leaving other areas designated as sensitive information unanonymized. This results in low accuracy and fails to fundamentally guarantee privacy. For example, existing videos allow for the inference of the photographer's location based on the background, thus reducing user privacy. Furthermore, existing scene reconstruction techniques primarily aim to accurately recreate the scene's appearance (shape, color, etc.), visually making the reconstructed scene as identical to the real scene as possible. However, directly applying this to vehicle video anonymization cannot guarantee the protection of sensitive information. This application addresses video anonymization based on scene event reconstruction. Scene events refer to events occurring within the scene, ignoring visual details (i.e., the three-dimensional shapes and colors in the scene are merely illustrative and cannot be directly visually correlated with the original scene). If existing scene reconstruction techniques are like "taking a photograph" of a scene, this invention is like drawing a "schematic diagram" of the scene. Compared to existing scene reconstruction technologies, this application uses different deep learning models and scene reconstruction logic to restore scene events (while the scene appearance uses a preset fixed 3D model). Unlike general desensitization algorithms that remove preset sensitive information from the original image, this application uses "reverse thinking" to add necessary information to be presented to the user on the basis of scene event reconstruction, thereby completely avoiding the leakage of sensitive information.
[0027] This application discloses a video anonymization method. In a system comprising a vehicle terminal, a cloud server, and a user terminal, the cloud server receives image features uploaded by the vehicle terminal. These image features include in-vehicle and out-of-vehicle image features. The image features are detected using a preset detection model to obtain detection information. Based on this detection information, a scene is constructed to obtain an anonymized reconstructed scene. The anonymized reconstructed scene is then sent to the user terminal. By using a detection model to detect the image features and obtain detection information, and then constructing a scene based on this detection information to obtain an anonymized reconstructed scene, this method avoids the inability to accurately anonymize all sensitive information (such as personal items and various account passwords) in a video. This video anonymization method uses a detection model to detect the image features, obtain detection information, and construct a scene based on this detection information to obtain an anonymized reconstructed scene. Furthermore, by using reverse thinking to collect necessary information from the image to construct the scene, the accuracy of video anonymization is improved.
[0028] Based on the above hardware structure, an embodiment of the video desensitization method of the present invention is proposed.
[0029] This invention provides a video desensitization method, referring to... Figure 2 , Figure 2This is a flowchart illustrating a first embodiment of a video desensitization method according to the present invention, wherein the video desensitization method is applied to a cloud server.
[0030] In this embodiment, the video desensitization method includes: Step S10: Receive image features uploaded by the vehicle terminal, wherein the image features include in-vehicle image features and out-of-vehicle image features; Step S20: Detect the image features based on a preset detection model to obtain detection information, and construct a scene based on the detection information to obtain a desensitized reconstruction scene.
[0031] In this embodiment, by processing the entire desensitization process on a cloud server, the processing pressure on the vehicle terminal can be reduced. This not only ensures the effective utilization of the vehicle terminal's internal processing resources but also improves the development and flexibility of the desensitization function by setting the main desensitization functions in the cloud. The vehicle terminal processes the original image to obtain image features, and the cloud server receives the image features uploaded by the vehicle terminal. Based on a preset detection model, the image features are detected to obtain detection information. Finally, based on the detection information, a scene is constructed to obtain a desensitized and reconstructed scene, thus completing the creation of the desensitized and reconstructed scene. The desensitized and reconstructed scene is then sent to the user terminal for the user to view. The image features include in-vehicle image features and out-of-vehicle image features. The detection model refers to a model for detecting image features, including at least a 3D human keypoint detection model and a 3D object detection model. By detecting the image features through the detection model, information such as the human body or object and their positional relationships are determined as detection information. Finally, based on the detection information, the scene is constructed to obtain the desensitized scene. The detection information refers to the detection position, shape, and orientation of the human body and object. The desensitized and reconstructed scene refers to a scene that does not contain sensitive content. In other words, the image features of the vehicle's interior / exterior are uploaded to the cloud. Pre-trained 3D human keypoint detection models and 3D object detection models are then used to detect these features, resulting in the classification and location of 3D keypoints of people and objects inside and outside the vehicle. This primarily utilizes detection models from visual perception (keypoint detection / object detection), unlike existing scene reconstruction techniques which directly predict the true shape and color of the scene. The pre-trained 3D human keypoint detection and 3D object detection models refer to models that have already been trained. Finally, the anonymized and reconstructed scene is built using image features. This approach avoids leaking privacy information, ensuring the accuracy and privacy of video anonymization. Furthermore, the use of visual perception detection models (keypoint detection / object detection), unlike existing scene reconstruction techniques which directly predict the true shape and color of the scene, reduces the cost of video anonymization and scene construction.
[0032] It's worth noting that for both in-vehicle and external scenes, multiple cameras can be used to extract images and process them to obtain image features. For 3D human keypoints or 3D objects detected by multiple cameras, the position and other information from the detection information used in scene reconstruction can be obtained by weighted averaging using preset weights. For example, if the position detected by camera 1 is X1 (3D position vector), and the position detected by camera 2 is X2 (3D position vector), then the final position X = w1 * X1 + w2 * X2, where w1>0, w2>0, and w1+w2=1 is the preset weight. This preset weight can be user-defined or set according to actual conditions. For example, users can choose to set a higher weight for the camera that is more accurate in actual use, or a higher weight for the camera that performs better in a particular environment. For instance, at night, a higher weight is set for the camera that performs better at night, and during the day, a higher weight is set for the camera that performs better at night. When only one camera detects a position, that position is used exclusively. The orientation, shape, and other information in the corresponding detection information can be determined according to the above preset weights, which can ensure the accuracy of the detection information determined by the features of the acquired image and further improve the accuracy of scene construction.
[0033] Furthermore, a schematic diagram of a video desensitization system is also provided for this embodiment, referring to... Figure 5In this embodiment, commonly used desensitization algorithms require running a detection model (backbone network + detection head) and performing image processing (decoding) on the vehicle side. This application, however, only requires running a backbone network on the vehicle side to encode images from in-vehicle / outside cameras using deep learning encoders 1 and 2, obtaining image features 1 and 2. Image feature 1 is then detected using a 3D human keypoint detection model and a 3D object detection model to obtain detection information related to humans and objects. Finally, in-vehicle scene reconstruction is performed based on the detection information. The processing flow for image feature 2 is the same as that for image feature 1, ultimately reconstructing the out-of-vehicle scene based on the detection information. After reconstructing the in-vehicle / outside scene, it is sent to the user terminal for display, thus realizing the display of 3D scene images on the user terminal. Users can select different angles to view the obtained 3D scene images on the terminal. The 3D human keypoint detection model, 3D object detection model, and deep learning encoder can all be used in-vehicle / outside environments; however, this increases the complexity of detection and encoding in each module and encoder, thereby reducing the space occupied by devices and model data. By deploying the anonymization computing resources on cloud servers, the vehicle terminal can perform only image acquisition, while the remaining steps are handled by the cloud server. This reduces the load on the vehicle terminal, and the remaining computations can be migrated to the cloud with its greater computing power. This approach offers fewer restrictions and greater flexibility in feature deployment and development. This application can quickly integrate existing functions, such as the sentinel mode: simply upload the key points detected in sentinel mode to the cloud (or migrate the key point detection head to the cloud to obtain key points), and the scene events can be reconstructed using the vehicle's 3D model, automatically completing the anonymization process. This significantly improves the accuracy of anonymization and reduces its cost.
[0034] This embodiment, comprising at least a vehicle terminal and a cloud server, acquires camera images via the vehicle terminal, including in-vehicle and external camera images. The camera images are encoded using a preset deep learning backbone network to obtain image features. These image features are uploaded to the cloud server, which then receives the image features uploaded from the vehicle terminal, including in-vehicle and external image features. The image features are detected using a preset detection model to obtain detection information, and a desensitized and reconstructed scene is constructed based on this detection information. By using a detection model to detect the image features and obtain detection information, and then constructing a scene based on this detection information to obtain a desensitized and reconstructed scene, the method avoids the inability to accurately desensitize all sensitive information in the video (such as personal items and various account passwords). This video desensitization method can detect the image features using a detection model, obtain detection information, and construct a scene based on this detection information to obtain a desensitized and reconstructed scene. Furthermore, by using reverse thinking to collect necessary information from the image to construct the scene, the accuracy of video desensitization is improved.
[0035] Furthermore, based on the first embodiment of the video desensitization method of the present invention, a second embodiment of the video desensitization method of the present invention is proposed, with reference to... Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the video desensitization method of this application, applied to a vehicle terminal. The video desensitization method includes: Step B10: Acquire camera image; Step B20: Encode the camera image based on a preset deep learning backbone network to obtain image features, wherein the image features include in-vehicle image features and out-of-vehicle image features; Step B30: Upload the image features to the cloud server, wherein the cloud server determines detection information based on the image features and constructs a de-identified reconstruction scene based on the detection information.
[0036] In this embodiment, images are acquired from cameras on the vehicle terminal, including images from both inside and outside the vehicle. These images are then encoded using a pre-trained deep learning backbone network (such as ResNet or ViT convolutional neural network) to obtain image features, i.e., inside / outside image features. These image features are represented as vectors in a high-dimensional image representation space (these features are determined by model training and cannot be extracted using methods other than the pre-trained model). Finally, the obtained image features are uploaded to a cloud server. The pre-trained deep learning backbone network can be divided into two parts: one for inside the vehicle and one for outside, processing the images from the inside and outside cameras respectively. Alternatively, the pre-trained deep learning backbone network can be located on the cloud server. The vehicle terminal only acquires camera images, ensuring the source of the anonymized images. The corresponding anonymization steps are then performed on the cloud server, specifically determining detection information based on image features and constructing an anonymized reconstruction scene based on that information. Encoding these features using a deep learning backbone network in the vehicle terminal provides a basis for subsequent scene reconstruction.
[0037] Furthermore, based on the second and first embodiments of the video desensitization method of the present invention, a third embodiment of the video desensitization method of the present invention is proposed, which is applied to a user terminal. The video desensitization method includes: Step C10: Obtain the input viewing angle switching command, and control the viewing angle of the desensitized and reconstructed scene based on the viewing angle switching command.
[0038] In this embodiment, on the user terminal, the viewing angle of the de-identified reconstructed scene is controlled based on the acquired viewing angle switching command. The viewing angle switching command refers to the command to switch the viewing angle of the de-identified reconstructed scene. This can be done by swiping on the user terminal or by directly inputting the viewing angle. For example, inputting a viewing angle of 30° east of north will display the resulting 3D de-identified reconstructed scene at an angle of 30° east of north. Because the de-identified reconstructed scene is constructed in 3D, the realism of the entire scene can be maintained at different angles, thereby improving the user experience.
[0039] Furthermore, based on the first, second, and third embodiments of the video desensitization method of the present invention, a fourth embodiment of the video desensitization method of the present invention is proposed, the video desensitization method comprising: Furthermore, the preset detection model includes an in-vehicle detection model and an out-of-vehicle detection model, and the detection information includes first detection information and second detection information. The step of detecting the image features based on the preset detection model to obtain detection information, and constructing a scene based on the detection information to obtain a desensitized and reconstructed scene includes: Step S21: Detect the features of the in-vehicle image based on the in-vehicle detection model to obtain first detection information, and construct a scene based on the first detection information to obtain a first desensitized reconstruction scene; Step S22: Detect the features of the vehicle exterior image based on the vehicle exterior detection model to obtain second detection information, and construct a second desensitized reconstruction scene based on the second detection information. Combine the first desensitized reconstruction scene and the second desensitized reconstruction scene as a desensitized reconstruction scene.
[0040] In this embodiment, it can be referred to Figure 5 When constructing the scene, two paths can be simultaneously performed: in-vehicle scene construction and out-of-vehicle scene construction. The in-vehicle detection model detects in-vehicle image features to obtain first detection information, and a first desensitized reconstruction scene is constructed based on this first detection information. Simultaneously, the out-of-vehicle detection model detects out-of-vehicle image features to obtain second detection information, and a second desensitized reconstruction scene is constructed based on this second detection information. Finally, the two desensitized reconstruction scenes are combined into a single desensitized reconstruction scene. Here, the first detection information refers to the detection results of in-vehicle image features, and the first desensitized reconstruction scene is the in-vehicle reconstruction scene; the second detection information refers to the detection results of out-of-vehicle image features, and the second desensitized reconstruction scene is the out-of-vehicle reconstruction scene. By performing scene reconstruction through both in-vehicle and out-of-vehicle paths simultaneously, the computational load of each model can be controlled, thereby improving the efficiency of scene reconstruction.
[0041] Furthermore, based on the first, second, third, and fourth embodiments of the video desensitization method of the present invention, a fifth embodiment of the video desensitization method of the present invention is proposed, the video desensitization method comprising: Furthermore, the preset detection model includes a 3D human keypoint detection model and a 3D object detection model. The step of detecting the image features based on the preset detection model to obtain detection information includes: Step a: Perform human detection on the image features based on the 3D human key point detection model to obtain human detection information; Step b: Perform object detection on the image features based on the 3D object detection model to obtain object detection information, and determine the detection information based on the human body detection information and the object detection information.
[0042] In this embodiment, the preset detection model includes at least a 3D human keypoint detection model and a 3D object detection model. After determining the image features, human or object detection is performed to obtain human detection information and object detection information. Human detection information refers to the information obtained by detecting image features based on the 3D human keypoint detection model, and object detection information refers to the information obtained by detecting image features based on the 3D object detection model. This detection information includes at least the human or object, as well as the required information for constructing the scene, such as the shape, size, position, or orientation of the human or object. A scene can then be constructed based on the detection information. Since the detection information only includes shape, size, position, or orientation, and does not directly reconstruct the scene based on the specific color or space occupancy of the image, the overall desensitization cost is kept low, and the process is simpler.
[0043] In this embodiment, human detection is performed on the image features based on a 3D human keypoint detection model to obtain human detection information, and object detection is performed on the image features based on a 3D object detection model to obtain object detection information. The human detection information and the object detection information are then used to determine the detection information. This detection information provides a basis for scene construction, and performing 3D detection based on the detection information ensures the stereoscopic effect of the constructed scene.
[0044] Furthermore, the detection model includes a clothing classification model and an age classification model. The step of performing human detection on the image features based on the 3D human keypoint detection model to obtain human detection information includes: Step c: Perform human body detection on the image features based on the 3D human key point detection model to determine the human image features; Step d: Based on the clothing classification model, feature detection is performed on the image features to obtain the first human body feature information, and based on the age classification model, feature detection is performed on the image features to obtain the second human body feature information; Step e: Based on the first human body feature information, the second human body feature information, and preset feature information, construct human body detection information by the human body image features, wherein the preset feature information includes the initial color features preset by the clothing classification model and the initial age group features preset by the age classification model.
[0045] In this embodiment, since the scene presented to the user is the reconstructed scene, it is theoretically impossible to directly obtain the original information of the in-vehicle / outside images, thus completely preventing the leakage of user privacy. Preferably, by adding a detection model, richer information can be presented to the user while fully ensuring privacy. After performing human body detection on the image features based on the 3D human body key point detection model and obtaining human body image features, the human body is further constructed. Here, human body image features refer to the relevant image features of the human body. Furthermore, a clothing classification model can be added to detect clothing in the reconstructed scene, using detected clothing as the first human feature information, such as wearing short sleeves. Because a classification model is used here, the reconstructed colors are preset, fixed colors, unlike existing scene reconstruction technologies that directly predict the actual colors of the scene. This ensures that the colors of clothing in the real scene and the colors in the reconstructed scene cannot be visually directly correlated; that is, the closest color is selected for construction based on the detected color, for example, only red is constructed without distinguishing between dark red and light red. A human age classification model can also be added, allowing age information to be displayed as the second human feature information in the scene reconstruction, such as children. Here, because a classification model is used, the reconstructed... The human body age is a preset fixed age range (and thus uses a preset fixed human body model corresponding to that age range), instead of inferring the real age of the human body in the scene like the models used in existing scene reconstruction technologies (directly predicting and outputting the human body model corresponding to the real age). This ensures that the human body in the real scene and the human body model in the reconstructed scene cannot be directly visually correlated. The first human body feature information refers to the color feature of the human body's clothing, and the second human body feature information refers to the feature of the human body's age, which enriches the human body construction. The preset initial color features and preset initial age range features of the clothing classification model mean that only broad divisions of color and age range are displayed, without refining color and age. For example, dark red and light red are both constructed as red, and 18 years old and 20 years old are both constructed as the age range of 15-35 years old. Through the clothing classification model and age classification model, the human body scene construction can be enriched, thereby improving the user experience. More models can also be used for detection according to actual needs or user requirements. Object detection can also be performed, and the realism of object construction can be improved by adding detection models.
[0046] Furthermore, the detection model also includes an animal 3D keypoint model, and the step of determining the detection information based on the human detection information and the object detection information includes: Step f: Based on the animal 3D key point model, perform animal detection on the image features to obtain animal detection information, and summarize the human detection information, the object detection information, and the animal detection information as detection information.
[0047] In this embodiment, when detecting humans and objects using a 3D human keypoint detection model and a 3D object detection model, to avoid missed or incorrect detection of animals, an animal 3D keypoint model is designed to detect animals in the image features, obtaining animal detection information. This animal detection information is then aggregated into the overall detection information, which includes animal-related information obtained from animal detection of image features. Finally, the human detection information, object detection information, and animal detection information are combined as the detection information for scene construction. By adding the animal 3D keypoint detection model, active targets other than humans, such as cats and dogs, can be displayed in scene reconstruction. Here, keypoint prediction is performed, and the reconstructed model uses a preset, fixed model, rather than directly predicting the shape / color of the target as used in existing scene reconstruction technologies. This ensures that animals in the real scene and the animal model in the reconstructed scene cannot be visually directly correlated. The use of the animal 3D keypoint model further improves the realism of the scene construction.
[0048] In this embodiment, feature detection is performed on the image features based on clothing classification models and age classification models to obtain first / second human body feature information. The human body image features are then updated based on the first / second human body feature information and preset feature information, or animal detection is performed on the image features based on an animal 3D keypoint model to obtain animal detection information. The human body detection information, the object detection information, and the animal detection information are then summarized as detection information. By adding detection models, the user experience can be guaranteed, and the visual effects of scene construction can be enriched on the basis of the original basic desensitization function.
[0049] Furthermore, the step of constructing a desensitized and reconstructed scene based on the detection information includes: Step g: Based on at least one of the object detection information, human body detection information, and animal detection information in the detection information and preset calibration coordinates, a desensitized and reconstructed scene is constructed.
[0050] In this embodiment, scene reconstruction based on detection information is divided into in-vehicle and out-of-vehicle scene reconstruction processes. By identifying object detection information, human detection information, and animal detection information contained in the detection information, and constructing object-desensitized scenes and human-desensitized scenes based on at least one of these and preset calibration coordinates, respectively, the scenes constructed using at least one detection information are summarized to obtain the human-desensitized scene. The calibration coordinates include in-vehicle and out-of-vehicle calibration coordinates, which can be pre-calibrated by the user according to actual conditions. For example, the pre-calibration of door positions, window positions, and seat positions can be used as reference points or reference objects for scene construction. In other words, when reconstructing the in-vehicle scene, object scene reconstruction is performed: the in-vehicle digital model is used to reconstruct the cabin interior scene, and through pre-calibrated coordinate transformation, the detected 3D objects (the pre-calibrated coordinates, which are the objects in the in-vehicle and out-of-vehicle calibration coordinates mentioned above) are placed into the interior scene according to the type of object detected, based on the position / size / orientation of the 3D object detected in the detection information; human scene reconstruction is performed: through pre-calibrated coordinate transformation, the 3D human key point model is placed into the interior scene according to the position and action represented by the detection information.
[0051] When reconstructing the exterior scene, object scene reconstruction is performed: the exterior scene is reconstructed using the vehicle's external model (for buildings other than the vehicle itself, this can be obtained by combining vehicle positioning and a 3D stereo map). Through pre-calibrated coordinate transformations, detected 3D objects (such as other vehicles) are placed into the exterior scene according to their detected position, size, and orientation based on their type. Human scene reconstruction is also performed: through pre-calibrated coordinate transformations, a pre-set 3D human model is placed into the exterior scene according to the position and actions indicated by detected 3D keypoints. The interior model refers to a pre-stored model of the vehicle's interior layout, which accurately determines the position of interior decorations and items. The exterior model refers to a pre-stored model of the vehicle's exterior layout. The pre-storage of both interior and exterior models ensures accurate scene construction. For example, the position of the seats can be accurately constructed based on the interior model, and the angle or size relationship of people or objects within the vehicle can be accurately constructed based on the exterior model, thus ensuring the accuracy of the constructed scene.
[0052] In this embodiment, object detection information is determined from the detection information, and a scene desensitization scene is constructed based on the object detection information and preset calibration coordinates. The calibration coordinates include in-vehicle and out-of-vehicle calibration coordinates. Human detection information is also determined from the detection information, and a human desensitization scene is constructed based on the human detection information and the calibration coordinates. The object desensitization scene and the human desensitization scene are then combined to obtain a desensitized reconstruction scene. By reconstructing desensitized scenes from inside / outside the vehicle, the accuracy of desensitization can be guaranteed. Furthermore, constructing scenes based on detection information significantly reduces the cost of scene construction compared to existing technologies.
[0053] Furthermore, the step of constructing a desensitized human scene based on the human detection information and the calibration coordinates includes: Step i: Determine all seat coordinates in the in-vehicle calibration coordinates and determine the in-vehicle position in the human body detection information; Step j: If the in-vehicle position does not match any of the seat coordinates, then determine the nearest seat coordinates corresponding to the in-vehicle position; Step k: Based on the nearest seat coordinates, construct a scene to obtain a human desensitization scene.
[0054] In this embodiment, when constructing the in-vehicle scene, due to limitations of the in-vehicle scene, key points of the lower body may be undetectable due to occlusion or other reasons. Therefore, for the position and movement of the lower body of the 3D human model, the nearest vehicle seat position and sitting posture are used by default. That is, by determining all seat coordinates in the in-vehicle calibration coordinate system and the in-vehicle position in the human body detection information, and then determining the nearest seat coordinate corresponding to the in-vehicle position in the in-vehicle calibration coordinate system when the in-vehicle position does not match any of the seat coordinates, the nearest seat coordinate corresponding to the in-vehicle position is determined. Finally, the scene is constructed based on the nearest seat coordinate to obtain the desensitized human body scene. Here, seat coordinates refer to the calibration coordinates of the seats in the in-vehicle calibration coordinate system, in-vehicle position refers to the coordinates of the human body inside the vehicle, and nearest seat coordinate refers to the coordinates of the nearest position to the human body's coordinates in the in-vehicle calibration coordinate system. This method can accurately determine the in-vehicle scene. When constructing the scene outside the vehicle, the window coordinates in the external calibration coordinates can be determined, and the internal position in the human body detection information can be determined accordingly. Then, based on the above steps, the human body desensitization scene outside the vehicle can be constructed. Alternatively, it can be constructed based on other positions, such as car doors, handles, wheels, etc.
[0055] In this embodiment, by determining all seat coordinates in the in-vehicle calibration coordinates and the in-vehicle position in the human body detection information, if the in-vehicle position does not match any of the seat coordinates, the nearest seat coordinate corresponding to the in-vehicle position is determined, and a human body desensitization scene is constructed based on the nearest seat coordinate. By constructing the scene based on reference objects in the existing digital model, the accuracy of the constructed scene can be guaranteed.
[0056] The present invention also provides a video desensitization system, referring to... Figure 3 The video desensitization system includes: Vehicle terminal A01 is used to acquire camera images; the camera images are encoded based on a preset deep learning backbone network to obtain image features, wherein the image features include in-vehicle image features and out-of-vehicle image features; the image features are uploaded to a cloud server, wherein the cloud server determines detection information based on the image features and constructs a de-identified reconstruction scene based on the detection information; Cloud server A02 is used to receive image features uploaded by vehicle terminals, wherein the image features include in-vehicle image features and out-of-vehicle image features; the image features are detected based on a preset detection model to obtain detection information, and a de-identified reconstructed scene is obtained based on the detection information.
[0057] Optionally, the vehicle terminal A01 is further configured to: Human detection is performed on the image features based on a 3D human key point detection model to obtain human detection information. Object detection is performed on the image features based on a 3D object detection model to obtain object detection information, and detection information is determined based on the human body detection information and the object detection information.
[0058] Optionally, the vehicle terminal A01 is further configured to: Human body detection is performed on the image features based on the 3D human key point detection model to obtain determined human image features. The image features are analyzed using a clothing classification model to obtain first human body feature information, and the image features are analyzed using an age classification model to obtain second human body feature information. Human detection information is obtained by constructing human image features based on the first human feature information, the second human feature information, and preset feature information. The preset feature information includes the initial color features preset by the clothing classification model and the initial age group features preset by the age classification model.
[0059] Optionally, the vehicle terminal A01 is further configured to: Animal detection is performed on the image features based on the animal 3D key point model to obtain animal detection information, and the human detection information, the object detection information and the animal detection information are summarized as detection information.
[0060] Optionally, the vehicle terminal A01 is further configured to: A desensitized and reconstructed scene is obtained by constructing a scene based on at least one of the object detection information, human body detection information, and animal detection information in the detection information and a preset calibration coordinate.
[0061] Optionally, the vehicle terminal A01 is further configured to: The in-vehicle image features are detected based on the in-vehicle detection model to obtain first detection information, and a first desensitized reconstruction scene is obtained by constructing a scene based on the first detection information. The vehicle exterior image features are detected based on the vehicle exterior detection model to obtain second detection information. A second desensitized reconstruction scene is then constructed based on the second detection information. The first desensitized reconstruction scene and the second desensitized reconstruction scene are then combined to form the desensitized reconstruction scene.
[0062] The methods executed by the above-mentioned program modules can be referred to in the various embodiments of the video desensitization method of the present invention, and will not be repeated here.
[0063] The present invention also provides a video desensitization device.
[0064] The device of the present invention includes: a memory, a processor, and a video desensitization program stored in the memory and executable on the processor. When the video desensitization program is executed by the processor, it implements the steps of the video desensitization method as described above.
[0065] The present invention also provides a storage medium.
[0066] The present invention stores a video desensitization program on a storage medium, which, when executed by a processor, implements the steps of the video desensitization method described above.
[0067] The method implemented when the video desensitization program running on the processor is executed can be referred to in various embodiments of the video desensitization method of the present invention, and will not be repeated here.
[0068] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0069] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0070] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0071] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A video anonymization method, characterized in that, The video desensitization method is applied to a cloud server, and the steps of the video desensitization method include: Receive image features uploaded by the vehicle terminal, wherein the image features include in-vehicle image features and out-of-vehicle image features; The image features are detected based on a preset detection model to obtain detection information, and a desensitized reconstructed scene is constructed based on the detection information. The preset detection model includes a 3D human keypoint detection model and a 3D object detection model. The step of detecting the image features based on the preset detection model to obtain detection information includes: Human detection is performed on the image features based on a 3D human keypoint detection model to obtain human detection information; object detection is performed on the image features based on a 3D object detection model to obtain object detection information, and detection information is determined based on the human detection information and the object detection information. The detection information includes 3D keypoints of the human body inside and outside the vehicle and the classification and position of objects inside and outside the vehicle. The step of constructing a desensitized and reconstructed scene based on the detection information includes: placing the detected 3D objects into the scene inside and outside the vehicle according to the detected object type based on the detected coordinate transformation; and placing the 3D human keypoint detection model into the scene inside and outside the vehicle according to the position and action represented by the detected 3D keypoints based on the pre-calibrated coordinate transformation.
2. The video desensitization method as described in claim 1, characterized in that, The step of performing human detection on the image features based on the 3D human key point detection model to obtain human detection information includes: Human image features are determined by performing human detection based on the image features using a 3D human key point detection model. The image features are analyzed using a clothing classification model to obtain first human body feature information, and the image features are analyzed using an age classification model to obtain second human body feature information. Human detection information is obtained by constructing human image features based on the first human feature information, the second human feature information, and preset feature information. The preset feature information includes the initial color features preset by the clothing classification model and the initial age group features preset by the age classification model.
3. The video desensitization method as described in claim 2, characterized in that, The detection model also includes an animal 3D key point model, and the step of determining the detection information based on the human body detection information and the object detection information includes: Animal detection is performed on the image features based on the animal 3D key point model to obtain animal detection information, and the human detection information, the object detection information and the animal detection information are summarized as detection information.
4. A video desensitization method, characterized in that, The video desensitization method is applied to a vehicle terminal, and the steps of the video desensitization method include: Acquire camera images; The camera images are encoded based on a pre-set deep learning backbone network to obtain image features, wherein the image features include in-vehicle image features and out-of-vehicle image features. The image features are uploaded to a cloud server, wherein the cloud server determines detection information based on the image features and constructs a desensitized reconstruction scene based on the detection information according to the method of claim 1.
5. A video desensitization system, characterized in that, The video desensitization system includes: The vehicle terminal is used to acquire camera images; the camera images are encoded based on a preset deep learning backbone network to obtain image features, wherein the image features include in-vehicle image features and out-of-vehicle image features; the image features are uploaded to a cloud server, wherein the cloud server determines detection information based on the image features and constructs a de-identified reconstruction scene based on the detection information; A cloud server is used to receive image features uploaded by vehicle terminals, wherein the image features include in-vehicle image features and out-of-vehicle image features; the image features are detected based on a preset detection model to obtain detection information, and a desensitized reconstructed scene is constructed based on the detection information, wherein the preset detection model includes a 3D human keypoint detection model and a 3D object detection model, and the step of detecting the image features based on the preset detection model to obtain detection information includes: Human detection is performed on the image features based on a 3D human keypoint detection model to obtain human detection information; object detection is performed on the image features based on a 3D object detection model to obtain object detection information, and detection information is determined based on the human detection information and the object detection information. The detection information includes 3D keypoints of the human body inside and outside the vehicle and the classification and position of objects inside and outside the vehicle. The step of constructing a desensitized reconstructed scene based on the detection information includes: placing the detected 3D objects into the scene inside and outside the vehicle according to the detected object type based on the detected coordinate transformation; and placing the 3D human keypoint detection model into the scene inside and outside the vehicle according to the position and action represented by the detected 3D keypoints based on the pre-calibrated coordinate transformation.
6. A video desensitization device, characterized in that, The video desensitization device includes: a memory, a processor, and a video desensitization program stored in the memory and executable on the processor, wherein the video desensitization program, when executed by the processor, implements the steps of the video desensitization method as described in any one of claims 1 to 4.
7. A computer storage medium, characterized in that, The computer storage medium stores a program for implementing the video desensitization method, which is executed by a processor to implement the steps of the video desensitization method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Scene reconstruction method and device, electronic equipment, program and medium
CN108230437A
Image desensitization method and related device
CN115408710A