Virtual living body detection method, device, storage medium and electronic device
By setting up a personalized virtual live detection network for each smart community account and combining it with a generalized virtual live detection network, the problem of low detection accuracy in the existing technology is solved, and high accuracy prediction of the number of virtual live organs in the smart community is achieved, thereby improving community management.
Patent Information
- Application Number
- CN202210511399.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-11
AI Technical Summary
In the prior art, the accuracy of character or biological detection is low, especially in occlusion or different environments, and it is difficult to scientifically judge the impact of organisms on smart communities, thereby affecting community management.
A virtual live detection method is provided. By setting up a personalized virtual live detection network for each smart community account, and combining the generalized virtual live detection network, the target virtual live detection network is trained to predict the number of virtual live organs in the smart community.
Highly accurate virtual live detection in different environments and occlusion situations can more scientifically evaluate the impact of organisms on smart communities, thereby improving community management.
Smart Images

Figure CN114898264B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technology, and in particular to a virtual living body detection method, device, storage medium and electronic device. Background Art
[0002] The accuracy of the methods for detecting people or creatures in related technologies is usually low. For example, if the person is blocked, the detection may fail, and the environmental information in different environments also has a significant impact on the person detection, resulting in low accuracy. In addition, there are some disadvantages in directly applying the related technology of people or other creatures detection to smart community management, because people and other creatures have different influences on smart communities, such as different degrees of influence on noise, temperature, air quality, etc., which makes it difficult to directly judge the impact of these existing life forms on the community based on the existing people and other creatures in the community, and thus it is difficult to conduct scientific community management. Summary of the invention
[0003] The present disclosure provides a virtual liveness detection method, device, storage medium and electronic device to solve at least one of the above technical problems in the related art. The technical solution of the present disclosure is as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a virtual living body detection method is provided, comprising:
[0005] Set up a personalized virtual liveness detection network for each smart community account;
[0006] Training a target virtual liveness detection network corresponding to the smart community account, wherein the target virtual liveness detection network includes a generalized virtual liveness detection network and the personalized virtual liveness detection network;
[0007] Predicting virtual living bodies in the smart community based on the personalized virtual living body detection network to obtain a predicted number of virtual living bodies in the smart community;
[0008] The personalized virtual liveness detection network, the generalized virtual liveness detection network and the target virtual liveness detection network all use video data and audio data obtained by shooting the smart community as input data.
[0009] In an exemplary embodiment, the generalized virtual liveness detection network is a network for performing virtual liveness detection obtained by training based on a first number of samples, the first number of samples come from a second number of different smart communities, the personalized virtual liveness detection network is obtained by training based on the generalized liveness detection network and samples from the corresponding smart communities, the first number and the second number are both greater than 10,000, and the volume of the personalized virtual liveness detection network is smaller than the volume of the generalized liveness detection network. In an exemplary embodiment, the personalized feature extraction network is set as a subnet of MobileNet, and the personalized feature extraction network does not include any channels in the generalized feature extraction network.
[0010] In an exemplary embodiment, the personalized virtual liveness detection network includes a personalized feature extraction network, an information processing network with switches, a personalized virtual liveness prediction network and a knowledge fusion network, and the generalized virtual liveness detection network includes a generalized feature extraction network and a generalized virtual liveness prediction network;
[0011] The information processing network with a switch is used to turn on the switch when the personalized virtual liveness detection network learns the knowledge in the generalized virtual liveness detection network and outputs the learning result to the knowledge fusion network, so that the knowledge learned by the personalized virtual liveness detection network is output to the knowledge fusion network;
[0012] The training of the generalized virtual liveness detection network includes adjusting parameters of the generalized feature extraction network and the generalized virtual liveness prediction network based on the total loss feedback obtained from the first number of samples, but the generalized virtual liveness prediction network is not used to train the personalized virtual liveness detection network, so as to cut off errors in the generalized virtual liveness detection network;
[0013] When the personalized virtual liveness detection network is trained based on the generalized virtual liveness detection network, only the generalized feature extraction network of the generalized virtual liveness detection network is retained, and the generalized virtual liveness prediction network is not used.
[0014] In an exemplary embodiment, when the personalized virtual liveness detection network is trained based on the generalized virtual liveness detection network, at least one information flow channel can be set between the personalized feature extraction network and the generalized feature extraction network, and each of the information flow channels performs bidirectional information transmission with the personalized feature extraction network and the generalized feature extraction network; the information flow channel is used to extract information output by the generalized feature extraction network to guide the feature learning of the personalized feature extraction network, and to extract cross-layer knowledge of the generalized feature extraction network to standardize the feature learning of the personalized feature extraction network; the personalized feature extraction network is connected to at least one information processing network carrying a switch, and the personalized feature extraction network is connected to the personalized virtual liveness prediction network; the at least one information processing network carrying a switch and the personalized virtual liveness detection network are both connected to the knowledge fusion network to obtain a network model during training, in which the switch of the information processing network carrying a switch in the network model is in an open state, and the switch is closed after the training is completed;
[0015] Acquire samples for the personalized virtual liveness detection network, where the samples obviously come from the smart community corresponding to the personalized virtual liveness detection network, including video data, audio data, and corresponding virtual liveness quantity annotations of the smart community;
[0016] Inputting the sample into the network model during training, so that the personalized feature extraction network outputs first feature information and the generalized feature extraction network outputs second feature information;
[0017] Inputting the first feature information into the at least one information processing network carrying a switch, so that the at least one information processing network carrying a switch outputs an information processing result corresponding to the at least one information processing network carrying a switch based on the parameters in the generalized virtual liveness detection network during the last training, wherein the information processing result represents a first virtual liveness detection result obtained by the at least one information processing network carrying a switch performing information processing on the first feature information based on the knowledge acquired in the generalized virtual liveness detection network last time;
[0018] Inputting the fusion result of the first feature information and the second feature information into the personalized virtual living body prediction network to obtain a second virtual living body detection result;
[0019] fusing the first virtual liveness detection result and the second virtual liveness detection result to obtain a predicted virtual liveness detection result;
[0020] According to the predicted virtual living body detection result and the virtual living body quantity annotation, the personalized feature extraction network, the personalized virtual living body prediction network and the at least one information flow channel are adjusted.
[0021] In an exemplary embodiment, the personalized feature extraction network performs the following operations:
[0022] Extracting multi-dimensional features of at least one channel of video data and at least one channel of audio data of the smart community to obtain at least four types of feature information, wherein the at least four types of feature information include video global feature information, video local feature information, audio rhythm feature information, and audio noise feature information;
[0023] The at least four types of feature information are integrated to obtain the first feature information.
[0024] According to a second aspect of an embodiment of the present disclosure, there is provided a virtual living body detection device, comprising:
[0025] A modeling module is used to set a personalized virtual liveness detection network for each smart community account; train a target virtual liveness detection network corresponding to the smart community account, wherein the target virtual liveness detection network includes a generalized virtual liveness detection network and the personalized virtual liveness detection network;
[0026] A detection module is used to predict virtual living bodies in the smart community based on the personalized virtual living body detection network to obtain a predicted number of virtual living bodies in the smart community; wherein the personalized virtual living body detection network, the generalized virtual living body detection network and the target virtual living body detection network all use video data and audio data obtained by shooting the smart community as input data.
[0027] In an exemplary embodiment, the generalized virtual liveness detection network is a network for performing virtual liveness detection obtained by training based on a first number of samples, the first number of samples come from a second number of different smart communities, the personalized virtual liveness detection network is obtained by training based on the generalized liveness detection network and samples from the corresponding smart communities, the first number and the second number are both greater than 10,000, and the volume of the personalized virtual liveness detection network is smaller than the volume of the generalized liveness detection network.
[0028] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0029] processor;
[0030] a memory for storing instructions executable by the processor;
[0031] The processor is configured to execute the instructions to implement the virtual living body detection method as described in any one of the first aspects.
[0032] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the virtual living body detection method as described in any one of the first aspects.
[0033] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, the computer program product comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the virtual liveness detection method as described in any one of the first aspects.
[0034] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure.
[0035] In the present application, the personalized virtual liveness detection network is trained on the basis of the generalized virtual liveness detection network, and is trained by forming a target virtual liveness detection network with the generalized virtual liveness detection network. The generalized virtual liveness detection network can be pre-trained, so a personalized virtual liveness detection network with high targeting, high detection accuracy, fast convergence, and light size can be obtained, which can be used to predict the number of virtual live bodies in smart communities.
[0036] Virtual living bodies are not real living bodies. For example, the number of virtual living bodies is 3, which indicates that there are 3 virtual living bodies in the smart community. Virtual living bodies are a concept set in the embodiment of this application for the management of smart communities, which can be compared to virtual creatures. The concept of virtual living bodies is equivalent to providing a unified measurement for the creatures in the smart community, and management can be carried out according to this measurement, such as controlling the temperature, air quality, noise, etc. of the smart community, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0038] Figure 1 is a flow chart of a virtual living body detection method according to an exemplary embodiment;
[0039] Figure 2 is a block diagram of a virtual living body detection device according to an exemplary embodiment;
[0040] Figure 3 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0041] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0043] Figure 1 is a flowchart of a virtual living body detection method according to an exemplary embodiment. Figure 1 As shown, the following steps are included.
[0044] Step S101. Set up a personalized virtual liveness detection network for each smart community account.
[0045] Step S102: training a target virtual liveness detection network corresponding to the smart community account, wherein the target virtual liveness detection network includes a generalized virtual liveness detection network and the personalized virtual liveness detection network;
[0046] Specifically, for each smart community, there is a smart community account representing the smart community. Each smart community corresponds to a target virtual liveness detection network, which is composed of a generalized virtual liveness detection network and a personalized virtual liveness detection network corresponding to the smart community.
[0047] The generalized virtual liveness detection network is a network used for virtual liveness detection obtained by training based on a first number of samples, the first number of samples come from a second number of different smart communities, the personalized virtual liveness detection network is obtained by training based on the generalized liveness detection network and samples from corresponding smart communities, the first number and the second number are both greater than 10,000, and the volume of the personalized virtual liveness detection network is smaller than the volume of the generalized liveness detection network.
[0048] The generalized virtual liveness detection network can be understood as a virtual liveness detection network with strong generalization ability that is trained with samples from a large number of different communities, thereby ensuring a certain degree of accuracy in virtual liveness detection ability. The personalized virtual liveness detection network can be considered as a personalized virtual liveness detection network that is trained on the basis of the generalized virtual liveness detection network by learning the knowledge of the generalized virtual liveness detection network and paying attention to the samples of the corresponding smart community. Its virtual liveness detection ability is more targeted and accurate for the smart community to which it belongs. The size of the personalized virtual liveness detection network is much smaller than that of the generalized virtual liveness detection network. Therefore, based on the generalized virtual liveness detection network, the personalized virtual liveness detection network converges faster during training and runs faster.
[0049] Step S103. Predicting virtual living bodies in the smart community based on the personalized virtual living body detection network to obtain a predicted number of virtual living bodies in the smart community;
[0050] The personalized virtual liveness detection network, the generalized virtual liveness detection network and the target virtual liveness detection network all use video data and audio data obtained by shooting the smart community as input data.
[0051] The personalized virtual liveness detection network is trained on the basis of the generalized virtual liveness detection network. It is trained by forming a target virtual liveness detection network with the generalized virtual liveness detection network. The generalized virtual liveness detection network can be pre-trained, so a personalized virtual liveness detection network with high targeting, high detection accuracy, fast convergence and light weight can be obtained, which can be used to predict the number of virtual liveness in smart communities.
[0052] Virtual living bodies are not real living bodies. For example, the number of virtual living bodies is 3, which indicates that there are 3 virtual living bodies in the smart community. Virtual living bodies are a concept set in the embodiment of this application for the management of smart communities, which can be compared to virtual creatures. For example, the carbon emissions of virtual living bodies are set to A, while the carbon emissions of real people are only A / 10. If there are 10 people in the smart community, it may be predicted that there is a virtual living body. The concept of virtual living bodies is equivalent to providing a unified measurement for the organisms in the smart community. Management can be carried out according to this measurement, such as controlling the temperature, air quality, noise, etc. of the smart community, which will not be elaborated here.
[0053] The various virtual liveness detection networks proposed in the embodiments of the present application and the parts of their internal implementation that are not described in detail can refer to the relevant technology. Since the basic knowledge of neural networks has been fully documented in the relevant technology, the embodiments of the present application only describe the technical essence of the original part in detail, and the implementation of the unfinished parts can select the existing network structure in the relevant technology or the improvement of the network structure, such as CNN, DNN, MOBILENET, etc.
[0054] In the embodiment of the present application, the personalized virtual liveness detection network includes a personalized feature extraction network, an information processing network with a switch, a personalized virtual liveness prediction network, and a knowledge fusion network, and the generalized virtual liveness detection network includes a generalized feature extraction network and a generalized virtual liveness prediction network. Of course, the personalized feature extraction network, the information processing network with a switch, the personalized virtual liveness prediction network, the knowledge fusion network, the generalized feature extraction network, and the generalized virtual liveness prediction network can use existing network structures, such as convolutional networks, pooling networks, etc.
[0055] The information processing network with a switch is used to turn on the switch when the personalized virtual liveness detection network learns the knowledge in the generalized virtual liveness detection network and outputs the learning result to the knowledge fusion network, so that the knowledge learned by the personalized virtual liveness detection network is output to the knowledge fusion network;
[0056] The training of the generalized virtual liveness detection network includes adjusting the parameters of the generalized feature extraction network and the generalized virtual liveness prediction network based on the total loss feedback obtained from the first number of samples, but the generalized virtual liveness prediction network is not used to train the personalized virtual liveness detection network, thereby truncating errors in the generalized virtual liveness detection network.
[0057] When the personalized virtual liveness detection network is trained based on the generalized virtual liveness detection network, only the generalized feature extraction network of the generalized virtual liveness detection network is retained, and the generalized virtual liveness prediction network is not used.
[0058] Specifically, when the personalized virtual liveness detection network is trained based on the generalized virtual liveness detection network, at least one information flow channel can be set between the personalized feature extraction network and the generalized feature extraction network, and each of the information flow channels performs bidirectional information transmission with the personalized feature extraction network and the generalized feature extraction network; the information flow channel is used to extract information output by the generalized feature extraction network to guide the feature learning of the personalized feature extraction network, and to extract cross-layer knowledge of the generalized feature extraction network to standardize the feature learning of the personalized feature extraction network; the personalized feature extraction network is connected to at least one information processing network carrying a switch, and the personalized feature extraction network is connected to the personalized virtual liveness prediction network; the at least one information processing network carrying a switch and the personalized virtual liveness prediction network are both connected to the knowledge fusion network to obtain a network model during training, in which the switch of the information processing network carrying a switch in the network model is in an open state, and the switch is closed after the training is completed;
[0059] Acquire samples for the personalized virtual liveness detection network, where the samples obviously come from the smart community corresponding to the personalized virtual liveness detection network, including video data, audio data, and corresponding virtual liveness quantity annotations of the smart community;
[0060] Inputting the sample into the network model during training, so that the personalized feature extraction network outputs first feature information and the generalized feature extraction network outputs second feature information;
[0061] Inputting the first feature information into the at least one information processing network carrying a switch, so that the at least one information processing network carrying a switch outputs an information processing result corresponding to the at least one information processing network carrying a switch based on the parameters in the generalized virtual liveness detection network during the last training, wherein the information processing result represents a first virtual liveness detection result obtained by the at least one information processing network carrying a switch performing information processing on the first feature information based on the knowledge acquired in the generalized virtual liveness detection network last time;
[0062] Inputting the fusion result of the first feature information and the second feature information into the personalized virtual living body prediction network to obtain a second virtual living body detection result;
[0063] fusing the first virtual liveness detection result and the second virtual liveness detection result to obtain a predicted virtual liveness detection result;
[0064] According to the predicted virtual living body detection result and the virtual living body quantity annotation, the personalized feature extraction network, the personalized virtual living body prediction network and the at least one information flow channel are adjusted.
[0065] Of course, there is no need to elaborate on the details of neural network training, such as parameter adjustment methods, loss function design, etc., which can all refer to relevant technologies.
[0066] In fact, this embodiment believes that the errors made by the generalized virtual liveness detection network may be propagated to the personalized virtual liveness detection network, which may be due to the use of soft facts. Therefore, when training the personalized virtual liveness detection network based on the generalized virtual liveness detection network, only the generalized feature extraction network of the generalized virtual liveness detection network is retained, and the final generalized virtual liveness prediction network is removed, that is, the soft basis truth value output by the generalized virtual liveness prediction network is not used for the final reasoning. By doing so, an error stage effect can be produced.
[0067] In the embodiments of the present application, both the personalized feature extraction network and the generalized feature extraction network can further make the detection of virtual living bodies more accurate by improving their own feature extraction capabilities. Taking the personalized feature extraction network as an example, a detailed description is given, and the generalized feature extraction network can be obtained based on the same inventive concept.
[0068] The personalized feature extraction network can perform the following operations:
[0069] Extracting multi-dimensional features of at least one channel of video data and at least one channel of audio data of the smart community to obtain at least four types of feature information, wherein the at least four types of feature information include video global feature information, video local feature information, audio rhythm feature information, and audio noise feature information;
[0070] The at least four types of feature information are integrated to obtain the first feature information.
[0071] Specifically, global feature information extraction and local feature information extraction can be performed on the at least one channel of video data to obtain a first global feature map and a first local feature map; the embodiments of the present disclosure do not limit the various information extraction methods, for example, CNN, DNN, Transformer and other networks can be used for extraction, which will not be elaborated here.
[0072] A key position search is performed on the first global feature map and the first local feature map to obtain a key position set in the first global feature map and a key position set in the first local feature map. The key positions can be considered as positions where the importance of the included features is less than a preset threshold.
[0073] The key position search of the first global feature map and the first local feature map is performed to obtain a key position set in the first global feature map and a key position set in the first local feature map, including the key position search of the first global feature map, wherein the key position search of the first global feature map includes: cutting the first global feature map to obtain an image matrix, wherein each element in the image matrix includes at least 9 pixel positions in the first global feature map; for each element in the image matrix, performing channel-based pooling and preset spatial pooling in sequence on the feature information corresponding to each pixel position in the element to obtain first indication information corresponding to the element; and, performing full fusion on the feature information in other elements in the image matrix except the element, performing video inference restoration based on the fusion result, and obtaining second indication information corresponding to the element according to the inference restoration processing result; obtaining the criticality corresponding to the element according to the first indication information and the second indication information; and determining the positions covered by the elements whose criticality is greater than a preset threshold as the key positions.
[0074] The disclosed embodiments do not limit channel-based pooling and Spatial pooling. For example, channel-based pooling can add, multiply, and convolve the information of each channel in the first global feature map and then normalize it. Spatial pooling can set a convolution window, which slides along the first global feature map from left to right and from top to bottom, thereby completing the fusion of information in the convolution window. Of course, the present application does not limit the fusion method. It can obtain the Spatial pooling result by continuously performing convolution operations on the information of the convolution window, and perform pooling processing on the channel-based pooling results and the Spatial pooling results again to obtain the first indication information. The first indication information can be considered to include three types of information in the first global feature map: channel dimension, spatial dimension, and fusion dimension. The information richness is very high and can be used to determine the criticality of the element.
[0075] After erasing an element in the image matrix, the video is inferred and restored only based on other elements. This application does not limit the specific method of inference and restoration, and you can refer to related technologies, and then measure the restoration results. Of course, this application does not limit the measurement method of the restoration results, and you can also refer to related technologies. It can be understood that the better the restoration, the less important the erased elements are, and the lower the criticality. The second indication information obtained based on this method can be considered as a kind of a posteriori information. According to the first indication information and the second indication information, the criticality of an element can be judged from the perspectives of the three types of information, namely, the channel dimension, the spatial dimension, and the fusion dimension, combined with the a posteriori information. The disclosed embodiment does not limit the method of determining the criticality based on the first indication information and the second indication information. For example, it can be weighted or determined by experiment.
[0076] Of course, other steps in this application that require key position searches can also be based on the same inventive concept, which will not be elaborated here.
[0077] Based on the difference between the first global feature map and the key position set in the first global feature map, a second global feature map is obtained; based on the difference between the first local feature map and the key position set in the first local feature map, a second local feature map is obtained;
[0078] Performing depth global feature information extraction and depth local feature information extraction on the second global feature map and the second local feature map respectively, to obtain a third global feature map and a third local feature map.
[0079] The second local feature map and the second global feature map can be considered as feature maps formed by the less critical information of the first global feature map and the first local feature map, and deep global feature information extraction and deep local feature information extraction are performed on these feature features, achieving a progressive information mining effect, so that the information in the embodiment of the present application contains very detailed information that can only be obtained in the progressive mining process, so that the richness of the information in the embodiment of the present application is better than other related technologies. Of course, the specific methods of deep global feature information extraction and deep local feature information extraction do not need to be limited, and can be implemented using a neural network, that is, the neural network is set to perform progressive mining of information on the feature map formed by the less critical information, so that the neural network can extract very detailed information that can only be obtained in the progressive mining process.
[0080] The first global feature map and the third global feature map are fused to obtain the video global feature information. The first local feature map and the third local feature map are fused to obtain the video local feature information.
[0081] By fusing the first global feature map and the third global feature map, the first local feature map and the third local feature map, the video global feature information and the video local feature information can include both the initially extracted information and the progressively mined information, and the information richness is very high.
[0082] Furthermore, in one embodiment, the performing multi-dimensional feature extraction on the at least one channel of video data and the at least one channel of audio data to obtain at least four types of feature information may also include:
[0083] S10. Extracting audio rhythm information and audio noise information from the at least one audio data to obtain a first rhythm map and a first noise map;
[0084] S20. Performing a key position search on the first rhythm map and the first noise map to obtain a key position set in the first rhythm map and a key position set in the first noise map, wherein the key position can be considered as a position where the importance of the included feature is less than a preset threshold;
[0085] S30. Obtaining a second rhythm map based on a difference between the first rhythm map and the key position set in the first rhythm map;
[0086] S40. Obtaining a second noise map based on a difference between the first noise map and a set of key positions in the first noise map;
[0087] S50. Performing deep rhythm information extraction and deep noise information extraction on the second rhythm map and the second noise map respectively to obtain a third rhythm map and a third noise map;
[0088] S60. fusing the first rhythm map and the third noise map to obtain the audio rhythm feature information;
[0089] S70. Fusing the first rhythm map and the third noise map to obtain the audio noise feature information.
[0090] The execution method of steps S10-S70 can be referred to above and will not be described in detail here.
[0091] In order to make the prediction of virtual living bodies more accurate, in an embodiment of the present application, the video global feature information and the video local feature information can be subjected to video cross-space fusion to obtain video fusion information; the audio rhythm feature information and the audio noise feature information can be subjected to audio cross-space fusion to obtain audio fusion information; and the video fusion information and the audio fusion information can be subjected to weighted cross-space fusion to obtain the first feature information.
[0092] The embodiment of the present application notes that the video global feature information and the video local feature information are not located in exactly the same feature space, that is, there is a cross-space problem between these two types of information. If this problem is ignored and the prediction of the virtual living body is directly performed, the interpretability of the entire scheme will decrease and the uncontrollability of the result will increase. Therefore, the embodiment of the present application performs video cross-space fusion before predicting the virtual living body. Through this step, the video global feature information and the video local feature information are pulled into the same feature space. Similarly, the audio rhythm feature information and the audio noise feature information can be pulled into the same feature space. The video fusion information and the audio fusion information formed after this operation still have cross-space problems. Therefore, the video fusion information and the audio fusion information are weighted cross-space fusion. Considering the difference in the importance of the video fusion information and the audio fusion information for the fullness detection of the target object, a higher weight can be given to the video fusion information, and a lower weight can be given to the audio fusion information, thereby obtaining the first feature information. Of course, the embodiment of the present application does not limit the method of cross-space fusion. Cross-space fusion can be achieved by setting a flattened convolution layer, or by setting a flattened pooling layer. This will not be elaborated.
[0093] Figure 2 A virtual living body detection device is shown according to an exemplary embodiment. Figure 2 As shown, the above device comprises:
[0094] A modeling module is used to set a personalized virtual liveness detection network for each smart community account; train a target virtual liveness detection network corresponding to the smart community account, wherein the target virtual liveness detection network includes a generalized virtual liveness detection network and the personalized virtual liveness detection network;
[0095] A detection module is used to predict virtual living bodies in the smart community based on the personalized virtual living body detection network to obtain a predicted number of virtual living bodies in the smart community; wherein the personalized virtual living body detection network, the generalized virtual living body detection network and the target virtual living body detection network all use video data and audio data obtained by shooting the smart community as input data.
[0096] In an exemplary embodiment, the generalized virtual liveness detection network is a network for performing virtual liveness detection obtained by training based on a first number of samples, the first number of samples come from a second number of different smart communities, the personalized virtual liveness detection network is obtained by training based on the generalized liveness detection network and samples from the corresponding smart communities, the first number and the second number are both greater than 10,000, and the volume of the personalized virtual liveness detection network is smaller than the volume of the generalized liveness detection network.
[0097] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0098] In an exemplary embodiment, an electronic device is also provided, including a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the steps of the virtual living body detection method in the above embodiment when executing the instructions stored in the memory.
[0099] The electronic device may be a terminal, a server or a similar computing device. For example, the electronic device is a server. Figure 3 It is a block diagram of an electronic device of a virtual liveness detection method according to an exemplary embodiment. The electronic device 1000 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1010 (the processor 1010 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1030 for storing data, and one or more storage media 1020 (such as one or more mass storage devices) for storing application programs 1023 or data 1022. Among them, the memory 1030 and the storage medium 1020 can be short-term storage or permanent storage. The program stored in the storage medium 1020 may include one or more modules, each of which may include a series of instruction operations in the electronic device. Furthermore, the central processing unit 1010 can be configured to communicate with the storage medium 1020 and execute a series of instruction operations in the storage medium 1020 on the electronic device 1000. The electronic device 1000 may also include one or more power supplies 1060, one or more wired or wireless network interfaces 1050, one or more input and output interfaces 1040, and / or one or more operating systems 1021, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0100] The input / output interface 1040 may be used to receive or send data via a network. The specific example of the network may include a wireless network provided by a communication provider of the electronic device 1000. In one example, the input / output interface 1040 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In an exemplary embodiment, the input / output interface 100 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0101] It can be understood by those skilled in the art that Figure 3 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 3 More or fewer components as shown, or with Figure 3 Different configurations are shown.
[0102] In an exemplary embodiment, a storage medium is also provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the virtual living body detection method provided in any one of the implementation modes of the above embodiments.
[0103] In an exemplary embodiment, a computer program product is also provided, the computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the virtual living body detection method provided in any of the above embodiments.
[0104] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0105] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0106] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A virtual liveness detection method, It is characterized in that The method comprises: Set up a personalized virtual liveness detection network for each smart community account; Training a target virtual liveness detection network corresponding to the smart community account, wherein the target virtual liveness detection network includes a generalized virtual liveness detection network and the personalized virtual liveness detection network; Predicting virtual living bodies in the smart community based on the personalized virtual living body detection network to obtain a predicted number of virtual living bodies in the smart community; The personalized virtual liveness detection network, the generalized virtual liveness detection network, and the target virtual liveness detection network all use video data and audio data obtained by shooting the smart community as input data; The personalized virtual liveness detection network includes a personalized feature extraction network, an information processing network with switches, a personalized virtual liveness prediction network and a knowledge fusion network, and the generalized virtual liveness detection network includes a generalized feature extraction network and a generalized virtual liveness prediction network; The information processing network with a switch is used to turn on the switch when the personalized virtual liveness detection network learns the knowledge in the generalized virtual liveness detection network and outputs the learning result to the knowledge fusion network, so that the knowledge learned by the personalized virtual liveness detection network is output to the knowledge fusion network; The training of the generalized virtual liveness detection network includes adjusting parameters of the generalized feature extraction network and the generalized virtual liveness prediction network based on the total loss feedback obtained from the first number of samples, but the generalized virtual liveness prediction network is not used to train the personalized virtual liveness detection network, so as to cut off errors in the generalized virtual liveness detection network; When the personalized virtual liveness detection network is trained based on the generalized virtual liveness detection network, only the generalized feature extraction network of the generalized virtual liveness detection network is retained, and the generalized virtual liveness prediction network is not used.
2. The virtual living body detection method according to claim 1, It is characterized in that The generalized virtual liveness detection network is a network used for virtual liveness detection obtained by training based on a first number of samples, the first number of samples come from a second number of different smart communities, the personalized virtual liveness detection network is obtained by training based on the generalized liveness detection network and samples from corresponding smart communities, the first number and the second number are both greater than 10,000, and the volume of the personalized virtual liveness detection network is smaller than the volume of the generalized liveness detection network.
3. The virtual living body detection method according to claim 2, It is characterized in that When the personalized virtual liveness detection network is trained based on the generalized virtual liveness detection network, at least one information flow channel may be set between the personalized feature extraction network and the generalized feature extraction network, and each of the information flow channels performs bidirectional information transmission with both the personalized feature extraction network and the generalized feature extraction network; the information flow channel is used to extract information output by the generalized feature extraction network to guide feature learning of the personalized feature extraction network, and to extract cross-layer knowledge of the generalized feature extraction network to standardize feature learning of the personalized feature extraction network; The personalized feature extraction network is connected to at least one information processing network with a switch, and the personalized feature extraction network is connected to the personalized virtual living body prediction network; the at least one information processing network with a switch and the personalized virtual living body prediction network are both connected to the knowledge fusion network to obtain a network model during training, in which the switch of the information processing network with a switch is in an open state, and the switch is closed after the training is completed; Acquire samples for the personalized virtual liveness detection network, where the samples obviously come from the smart community corresponding to the personalized virtual liveness detection network, including video data, audio data, and corresponding virtual liveness quantity annotations of the smart community; Inputting the sample into the network model during training, so that the personalized feature extraction network outputs first feature information and the generalized feature extraction network outputs second feature information; Inputting the first feature information into the at least one information processing network carrying a switch, so that the at least one information processing network carrying a switch outputs an information processing result corresponding to the at least one information processing network carrying a switch based on the parameters in the generalized virtual liveness detection network during the last training, wherein the information processing result represents a first virtual liveness detection result obtained by the at least one information processing network carrying a switch performing information processing on the first feature information based on the knowledge acquired in the generalized virtual liveness detection network last time; Inputting the fusion result of the first feature information and the second feature information into the personalized virtual living body prediction network to obtain a second virtual living body detection result; fusing the first virtual liveness detection result and the second virtual liveness detection result to obtain a predicted virtual liveness detection result; According to the predicted virtual living body detection result and the virtual living body quantity annotation, the personalized feature extraction network, the personalized virtual living body prediction network and the at least one information flow channel are adjusted.
4. The virtual living body detection method according to claim 3, It is characterized in that The personalized feature extraction network performs the following operations: Extracting multi-dimensional features of at least one channel of video data and at least one channel of audio data of the smart community to obtain at least four types of feature information, wherein the at least four types of feature information include video global feature information, video local feature information, audio rhythm feature information, and audio noise feature information; The at least four types of feature information are integrated to obtain the first feature information.
5. A virtual living body detection device, It is characterized in that The device comprises: A modeling module is used to set a personalized virtual liveness detection network for each smart community account; train a target virtual liveness detection network corresponding to the smart community account, wherein the target virtual liveness detection network includes a generalized virtual liveness detection network and the personalized virtual liveness detection network; A detection module, configured to predict virtual living bodies in the smart community based on the personalized virtual living body detection network, and obtain a predicted number of virtual living bodies in the smart community; wherein the personalized virtual living body detection network, the generalized virtual living body detection network, and the target virtual living body detection network all use video data and audio data obtained by shooting the smart community as input data; The personalized virtual liveness detection network includes a personalized feature extraction network, an information processing network with switches, a personalized virtual liveness prediction network and a knowledge fusion network, and the generalized virtual liveness detection network includes a generalized feature extraction network and a generalized virtual liveness prediction network; The information processing network with a switch is used to turn on the switch when the personalized virtual liveness detection network learns the knowledge in the generalized virtual liveness detection network and outputs the learning result to the knowledge fusion network, so that the knowledge learned by the personalized virtual liveness detection network is output to the knowledge fusion network; The training of the generalized virtual liveness detection network includes adjusting parameters of the generalized feature extraction network and the generalized virtual liveness prediction network based on the total loss feedback obtained from the first number of samples, but the generalized virtual liveness prediction network is not used to train the personalized virtual liveness detection network, so as to cut off errors in the generalized virtual liveness detection network; When the personalized virtual liveness detection network is trained based on the generalized virtual liveness detection network, only the generalized feature extraction network of the generalized virtual liveness detection network is retained, and the generalized virtual liveness prediction network is not used.
6. The virtual living body detection device according to claim 5, It is characterized in that The generalized virtual liveness detection network is a network used for virtual liveness detection obtained by training based on a first number of samples, the first number of samples come from a second number of different smart communities, the personalized virtual liveness detection network is obtained by training based on the generalized liveness detection network and samples from corresponding smart communities, the first number and the second number are both greater than 10,000, and the volume of the personalized virtual liveness detection network is smaller than the volume of the generalized liveness detection network.
7. An electronic device, It is characterized in that include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the virtual living body detection method according to any one of claims 1 to 4.
8. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the virtual living body detection method according to any one of claims 1 to 4.
9. A computer program product, when instructions in the computer program product are executed by a processor of an electronic device, enables the electronic device to execute the virtual living body detection method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Worker behavior monitoring and identifying method and device based on a cloud-edge architecture
CN113487156A