Method and device for determining audience number based on hybrid perception
By combining visual perception and signal perception methods to comprehensively determine the number of viewers, the problem of inaccurate determination of the number of viewers in the existing technology is solved, and more efficient audience number calculation and data support for subsequent analysis are achieved.
Patent Information
- Application Number
- CN202111500513.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2041-12-09
AI Technical Summary
When determining the number of viewers, existing technologies usually only analyze through a single perception dimension, resulting in inaccurate and inefficient determination of the number of viewers, and unable to meet the precise requirements of indicators such as the number of people reached by advertisements.
Combining the two methods of visual perception and signal perception, the number of visually perceived viewers and the number of signal-perceived viewers are determined respectively through image recognition algorithm and signal positioning algorithm, and comprehensive corrections are made on this basis to obtain the number of viewers in the target area.
It achieves more accurate and efficient determination of the number of viewers, provides a real data basis for subsequent viewing indicator analysis, and improves the accuracy and efficiency of determining the number of viewers.
Smart Images

Figure CN114241555B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audience number calculation, and in particular to a method and device for determining audience number based on hybrid perception. Background Art
[0002] Nowadays, various data analysis entities are increasingly focusing on the number of times a specific media is viewed by viewers to determine how many viewers a specific content has reached. For example, when placing advertisements, advertisers are increasingly concerned with how many people the ads reach, rather than how many devices they are played on. Existing technologies generally determine this metric by analyzing the number of viewers within the playback area. However, these techniques typically determine the number of viewers based on a single perceptual dimension, such as image recognition analysis algorithms, without considering the comprehensive determination of the number of viewers by combining data from multiple perceptual dimensions. Therefore, a more accurate and efficient method for determining the number of viewers is urgently needed. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method and device for determining the number of viewers based on hybrid perception, which can determine the number of viewers more accurately and efficiently, and provide a real data basis for subsequent viewing indicator analysis.
[0004] In order to solve the above technical problems, the first aspect of the present invention discloses a method for determining the number of viewers based on hybrid perception, the method comprising:
[0005] Acquiring audience image information of a target area, and determining the number of visually perceived audience members in the target area based on the audience image information;
[0006] Acquiring device communication signals in the target area, and determining the number of signal-perceiving viewers in the target area based on the device communication signals;
[0007] The number of viewers in the target area is determined according to the number of visually perceived viewers and the number of signal-perceived viewers.
[0008] As an optional embodiment, in the first aspect of the present invention, determining the number of visually perceived viewers in the target area based on the audience image information includes:
[0009] Determining facial features in the audience image information based on an image recognition algorithm;
[0010] The number of visually perceived viewers in the target area is determined based on the facial features.
[0011] As an optional embodiment, in the first aspect of the present invention, determining the number of visually perceived viewers in the target area based on the facial features includes:
[0012] For any of the facial features, calculating the viewing intention of the facial feature;
[0013] determining whether the viewing intention is within a preset viewing intention score range corresponding to the target area, and if the judgment result is yes, determining the facial feature as a viewing facial feature;
[0014] The total number of all viewed facial features is determined as the number of visually perceived viewers in the target area.
[0015] As an optional embodiment, in the first aspect of the present invention, the device communication signal includes wireless communication signals of multiple terminal devices; and determining the number of signal-perceiving viewers in the target area based on the device communication signal includes:
[0016] Determine the location information of each terminal device based on the wireless communication signals of the plurality of terminal devices based on a signal positioning algorithm;
[0017] For any of the terminal devices, determining whether the location information is within the target area, and if the determination result is yes, determining that the terminal device is a terminal within the area;
[0018] The total number of terminals in all the areas is determined as the number of signal-perceiving viewers in the target area.
[0019] As an optional embodiment, in the first aspect of the present invention, determining the number of viewers in the target area based on the number of visually perceived viewers and the number of signal-perceived viewers includes:
[0020] Determining whether the number of visually perceived viewers is consistent with the number of signal-perceived viewers;
[0021] When the judgment result is consistent, determining the number of visually perceived viewers or the number of signal-perceived viewers as the number of viewers in the target area;
[0022] When the judgment result is inconsistent, the visually perceived number of viewers is corrected to obtain the number of viewers in the target area.
[0023] As an optional embodiment, in the first aspect of the present invention, the correcting the number of visually perceived viewers to obtain the number of viewers in the target area includes:
[0024] Determining whether the number of visually perceived viewers is greater than the number of signal-perceived viewers, and if so, determining the number of visually perceived viewers as the number of viewers in the target area;
[0025] If not, calculating the difference between the number of visually perceived viewers and the number of signal-perceived viewers;
[0026] Determining whether the difference is greater than a preset difference threshold;
[0027] If yes, determining the number of visually perceived viewers as the number of viewers in the target area;
[0028] If not, the visually perceived number of viewers is corrected to obtain the number of viewers in the target area.
[0029] As an optional embodiment, in the first aspect of the present invention, the number of visually perceived viewers is calculated by a local device deployed in the target area based on an offline algorithm model; the number of signal-perceived viewers is calculated by a cloud device; and the correction of the number of visually perceived viewers to obtain the number of viewers within the target area includes:
[0030] Sending the audience image information to the cloud device, so that the cloud device determines the number of cloud visually perceived audiences corresponding to the audience image information based on an online image recognition algorithm model;
[0031] Determining whether the number of cloud-based visually perceived viewers is consistent with the number of visually perceived viewers;
[0032] If the judgment result is no, the number of viewers in the target area is determined according to the number of cloud-based visual perception viewers and the number of signal perception viewers.
[0033] A second aspect of an embodiment of the present invention discloses a device for determining the number of viewers based on hybrid perception, the device comprising:
[0034] A visual perception module, configured to obtain image information of audience members in a target area, and determine the number of visually perceived audience members in the target area based on the image information of the audience members;
[0035] a signal sensing module, configured to obtain device communication signals in the target area, and determine the number of signal sensing viewers in the target area based on the device communication signals;
[0036] The number determination module is used to determine the number of audience members in the target area according to the number of visually perceived audience members and the number of signal-perceived audience members.
[0037] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the visual perception module determines the number of visually perceived viewers in the target area based on the audience image information includes:
[0038] Determining facial features in the audience image information based on an image recognition algorithm;
[0039] The number of visually perceived viewers in the target area is determined based on the facial features.
[0040] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the visual perception module determines the number of visually perceived viewers in the target area based on the facial features includes:
[0041] For any of the facial features, calculating the viewing intention of the facial feature;
[0042] determining whether the viewing intention is within a preset viewing intention score range corresponding to the target area, and if the judgment result is yes, determining the facial feature as a viewing facial feature;
[0043] The total number of all viewed facial features is determined as the number of visually perceived viewers in the target area.
[0044] As an optional embodiment, in the second aspect of the present invention, the device communication signal includes wireless communication signals of multiple terminal devices; and the signal perception module determines the number of signal perception viewers in the target area based on the device communication signal in a specific manner, including:
[0045] Determine the location information of each terminal device based on the wireless communication signals of the plurality of terminal devices based on a signal positioning algorithm;
[0046] For any of the terminal devices, determining whether the location information is within the target area, and if the determination result is yes, determining that the terminal device is a terminal within the area;
[0047] The total number of terminals in all the areas is determined as the number of signal-perceiving viewers in the target area.
[0048] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the number determination module determines the number of viewers in the target area based on the number of visually perceived viewers and the number of signal-perceived viewers includes:
[0049] Determining whether the number of visually perceived viewers is consistent with the number of signal-perceived viewers;
[0050] When the judgment result is consistent, determining the number of visually perceived viewers or the number of signal-perceived viewers as the number of viewers in the target area;
[0051] When the judgment result is inconsistent, the visually perceived number of viewers is corrected to obtain the number of viewers in the target area.
[0052] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the number determination module corrects the visually perceived number of viewers to obtain the number of viewers in the target area includes:
[0053] Determining whether the number of visually perceived viewers is greater than the number of signal-perceived viewers, and if so, determining the number of visually perceived viewers as the number of viewers in the target area;
[0054] If not, calculating the difference between the number of visually perceived viewers and the number of signal-perceived viewers;
[0055] Determining whether the difference is greater than a preset difference threshold;
[0056] If yes, determining the number of visually perceived viewers as the number of viewers in the target area;
[0057] If not, the visually perceived number of viewers is corrected to obtain the number of viewers in the target area.
[0058] As an optional embodiment, in the second aspect of the present invention, the number of visually perceived viewers is calculated by the visual perception module triggering a local device deployed in the target area based on an offline algorithm model; the number of signal-perceived viewers is calculated by the signal perception module triggering a cloud device; and the specific manner in which the number determination module corrects the number of visually perceived viewers to obtain the number of viewers within the target area includes:
[0059] Sending the audience image information to the cloud device, so that the cloud device determines the number of cloud visually perceived audiences corresponding to the audience image information based on an online image recognition algorithm model;
[0060] Determining whether the number of cloud-based visually perceived viewers is consistent with the number of visually perceived viewers;
[0061] If the judgment result is no, the number of viewers in the target area is determined according to the number of cloud-based visual perception viewers and the number of signal perception viewers.
[0062] A third aspect of the present invention discloses another apparatus for determining the number of viewers based on hybrid perception, the apparatus comprising:
[0063] a memory storing executable program code;
[0064] a processor coupled to the memory;
[0065] The processor calls the executable program code stored in the memory to execute part or all of the steps in the method for determining the number of audience members based on hybrid perception disclosed in the first aspect of the present invention.
[0066] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0067] In an embodiment of the present invention, a method and device for determining the number of viewers based on hybrid perception are disclosed. The method includes: obtaining audience image information of a target area, and determining the number of visually perceived viewers in the target area based on the audience image information; obtaining device communication signals in the target area, and determining the number of signal-perceived viewers in the target area based on the device communication signals; and determining the number of viewers within the target area based on the number of visually perceived viewers and the number of signal-perceived viewers. It can be seen that the embodiment of the present invention can use the number of visually perceived viewers and the number of signal-perceived viewers to comprehensively determine the number of viewers in the area, thereby being able to determine the number of viewers more accurately and efficiently, providing a real data basis for subsequent viewing indicator analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0069] Figure 1 This is a flow chart of a method for determining the number of viewers based on hybrid perception disclosed in an embodiment of the present invention.
[0070] Figure 2 It is a structural diagram of a device for determining the number of viewers based on hybrid perception disclosed in an embodiment of the present invention.
[0071] Figure 3 It is a structural diagram of another apparatus for determining the number of viewers based on hybrid perception disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0072] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0073] The terms "first," "second," and so on, in the description and claims of the present invention and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or device.
[0074] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0075] The present invention discloses a method and device for determining the number of viewers based on hybrid perception. This method utilizes both visually perceived and signal-based audience numbers to comprehensively determine the number of viewers within a region, thereby enabling more accurate and efficient determination of the number of viewers, providing a reliable data foundation for subsequent viewing indicator analysis. These are described in detail below.
[0076] Example 1
[0077] See also Figure 1 , Figure 1 This is a flow chart of a method for determining the number of viewers based on hybrid perception disclosed in an embodiment of the present invention. Figure 1 The method for determining the number of viewers described can be applied to a viewer number calculation chip, a calculation terminal, or a calculation server (wherein the calculation server can be a local server or a cloud server). Figure 1 As shown, the method for determining the number of audience members based on hybrid perception may include the following operations:
[0078] 101. Obtain audience image information of a target area, and determine the number of visually perceived audience members in the target area based on the audience image information.
[0079] Optionally, the target area may be an indoor area or an outdoor area. The target area may be equipped with a playback device for playing specific media, such as an OTT TV (Over The Top TV) or a large advertising screen. The method described in the present invention can then be used to analyze the number of viewers of the played specific media. For example, the target area may be a living room in a home equipped with a smart TV.
[0080] Alternatively, the audience image information may be acquired by an image acquisition device, such as a camera, located in the target area. Alternatively, the image acquisition device may be located on the aforementioned playback device, such as an OTTTV integrated with a camera.
[0081] Optionally, the number of visually perceived viewers in the target area may be determined based on the audience image information using an image recognition algorithm.
[0082] 102. Obtain device communication signals in the target area, and determine the number of signal-perceiving viewers in the target area based on the device communication signals.
[0083] Optionally, the device communication signal may be a wired signal or a wireless signal. Optionally, the device communication signal may be a Bluetooth signal, a Wi-Fi signal, or other electromagnetic wave signal that can be used for communication. Optionally, the device communication signal may be received by a signal receiving device located in the target area and transmitted to a cloud server for analysis to determine the number of signal-aware viewers in the target area based on the device communication signal.
[0084] 103. Determine the number of audience members in the target area based on the number of visually perceived audience members and the number of signal-perceived audience members.
[0085] Optionally, the number of viewers in the target area may be determined by calculating the sum or difference between the number of visually perceived viewers and the number of signal-perceived viewers, or by calculating the number of visually perceived viewers and the number of signal-perceived viewers under specific rules.
[0086] Optionally, the number of audience members in the target area can be used to determine audience portrait information. For example, the portrait information of each audience member in the target area, such as age, occupation, hobbies, etc., can be determined as the audience portrait of the target area for subsequent data analysis.
[0087] It can be seen that the above-mentioned embodiments of the invention can use the number of viewers perceived by vision and the number of viewers perceived by signals to comprehensively determine the number of viewers in the area, thereby determining the number of viewers more accurately and efficiently, and providing a real data basis for subsequent viewing indicator analysis.
[0088] As an optional implementation, in step 101, determining the number of visually perceived viewers in the target area based on the viewer image information includes:
[0089] Determine facial features in the audience's image information based on image recognition algorithms;
[0090] Based on facial features, the number of visually perceived viewers in the target area is determined.
[0091] Optionally, the audience image information may be two-dimensional image information or three-dimensional image information, wherein the image recognition algorithm may be a two-dimensional face image recognition algorithm, such as a face template matching algorithm, an image matrix singular value feature recognition algorithm, a subspace analysis method, a local preserving projection algorithm, a principal component analysis method, an elastic matching method, a KL transform-based eigenface method, an artificial neural network method, a support vector machine method, an integral image feature method, or a probability model-based method. Optionally, the image recognition algorithm may also be a three-dimensional face image recognition algorithm, such as a 3D local matching method based on image features or a method based on model variable parameters, wherein the method based on model variable parameters combines the 3D deformation of a universal face model with matrix iterative minimization based on distance mapping to restore the head posture and 3D face, continuously updating the posture parameters as the correlation relationship of the model deformation changes, and repeating this process until the minimization scale meets the requirement.
[0092] Optionally, the number of visually perceived viewers in the target area may be determined based on the number of facial features in the target area.
[0093] It can be seen that by implementing this optional implementation method, the facial features in the audience image information can be determined based on the image recognition algorithm, and the number of visually perceived viewers in the target area can be determined based on the facial features, so that the number of viewers perceived in the visual dimension can be accurately determined, which is helpful for the subsequent accurate and comprehensive determination of the number of viewers in the area, and provides a real data basis for subsequent viewing indicator analysis.
[0094] As an optional implementation, in the above step, determining the number of visually perceived viewers in the target area based on facial features includes:
[0095] For any facial feature, calculate the viewing intention of the facial feature;
[0096] Determining whether the viewing intention is within a preset viewing intention score range corresponding to the target area, and if the determination result is yes, determining the facial feature as a viewing facial feature;
[0097] The total number of all viewed facial features is determined as the number of visually perceived viewers in the target area.
[0098] Optionally, viewing intention may include one or a combination of facial orientation angle and eye gaze angle. For example, viewing intention may be the facial orientation angle detected by a facial orientation angle detection regression model, or the eye gaze angle detected by an eye gaze angle detection regression model, or the user viewing score detected by a user viewing detection regression model that can simultaneously detect facial orientation angle and eye gaze angle. Accordingly, the preset viewing intention score range may be a preset facial orientation angle range, eye gaze angle range, or user viewing score range.
[0099] Optionally, a face orientation angle detection regression model can be trained. As an example, ResNet18 can be used as the network model skeleton of the face orientation angle detection regression model. Collect and annotate face orientation angle data, use face images containing pre-annotated face orientation angle values (each face image includes face orientation angle values in three directions: pitch, yaw, and roll) as the first training sample, and use the first training sample to train the face orientation angle detection regression model to obtain a trained face orientation angle detection regression model. Specifically, during training, the first training sample is used as the input of the face posture regression model, and the output is the face orientation angle values in three directions: pitch, yaw, and roll. When the average error of the output angles in the three directions is less than the preset angle (such as 5 degrees), stop training.
[0100] Optionally, a human eye gaze angle detection regression model can be trained. As an example, VGG19 can be used as the network model skeleton of the human eye gaze angle detection regression model. Collect and annotate human eye gaze angle data, and use human eye images containing pre-annotated human eye gaze angle values (each eye image includes the human eye gaze angle values of the left eye in the three directions of pitch, yaw, and roll and the annotated values of the human eye gaze angle values of the right eye in the three directions of pitch, yaw, and roll) as the second training sample, and use the second training sample to train the human eye gaze angle detection regression model to obtain a trained human eye gaze angle detection regression model. Specifically, the second training sample is used as the input of the eye gaze angle detection regression model, and the output is a set of eye gaze angles in three directions (left eye + right eye pitch, yaw, and roll). That is, the output of the eye gaze angle detection regression model is six angle values in total, namely the eye gaze angle values in the three directions of pitch, yaw, and roll of the left eye and the eye gaze angle values in the three directions of pitch, yaw, and roll of the right eye. When the average error of the output angles in the three directions is less than 8 degrees, the training is stopped. It should be noted that the eye image containing the pre-marked eye gaze angles can be a whole face image collected in advance (that is, the eye gaze angle values are pre-marked on the face image), or only the eye area part in the face image can be input into the gaze regression model as the second training sample. Those skilled in the art can select the corresponding training sample according to actual needs.
[0101] Optionally, a user viewing detection regression model can be trained. Specifically, a collection device is first used to collect multiple face videos, each video is about 2 seconds (a total of about 60 frames of images). The collected data is labeled and divided into two categories: images of viewed faces and images of unviewed faces, and the labeled face images are used as the third training sample. Then, the third training sample is input into the above-mentioned trained face orientation angle detection regression model and the trained eye gaze angle detection regression model to obtain face orientation angle values (angle values in the three directions of pitch, yaw, and roll) and eye gaze angle values (left eye gaze angle values in the three directions of pitch, yaw, and roll and right eye gaze angle values in the three directions of pitch, yaw, and roll). In this way, 9 angle values can be obtained for each frame of face image. As an example, the angle values of 10 frames of images can be randomly extracted from each face video data as the input of the user viewing detection regression model. In this way, 90 angle values can be obtained for each face video data (10 frames of face images are randomly extracted from each video data, and 9 angle values are obtained for each frame of face image).
[0102] Optionally, a multi-layer perceptron (MLP) can be used to train a user viewing detection regression model for classification. The 90 angle values obtained from each face video data set are used as input features for the user viewing detection regression model. The model is trained and the output is the user viewing score. The trained model is tested. If it fails the test, the model is retrained until it passes the test, completing the training.
[0103] Specifically, when calculating the viewing willingness of facial features, the facial features can be input into a facial orientation angle detection regression model or a human eye gaze angle detection regression model to obtain corresponding angle values, and the corresponding angle values are used as the viewing willingness. Alternatively, multiple frames of facial images related to the facial features can be input into a trained facial orientation angle detection regression model, and each frame of facial image obtains facial orientation angle values in the three directions of pitch, yaw, and roll. At the same time, multiple frames of facial images can be input into a trained human eye gaze angle detection regression model, and each frame of facial image obtains left eye gaze angle values in the three directions of pitch, yaw, and roll, and right eye gaze angle values in the three directions of pitch, yaw, and roll. Then, the obtained facial orientation angle values and human eye gaze angle values of the multiple frames of facial images (a total of 90 angle values as an input feature) are input into a trained user viewing detection regression model to obtain a user viewing score, and the user viewing score is used as the viewing willingness of the facial feature.
[0104] As an optional embodiment, the device communication signal includes wireless communication signals of multiple terminal devices; in the above step 102, determining the number of signal-perceiving viewers in the target area based on the device communication signal includes:
[0105] Based on the signal positioning algorithm, the location information of each terminal device is determined according to the wireless communication signals of multiple terminal devices;
[0106] For any terminal device, determine whether the location information is within the target area, and if the judgment result is yes, determine that the terminal device is a terminal within the area;
[0107] The total number of terminals in all areas is determined as the number of signal-perceiving viewers in the target area.
[0108] Optionally, the signal positioning algorithm may include at least one of a weight-based positioning algorithm, a fingerprint-based positioning algorithm, and a positioning method based on triangulation. Optionally, the calculation basis of the weight-based positioning algorithm includes at least one of a distance weight, a signal accuracy weight, and a signal source weight. Optionally, the positioning strategy of the fingerprint-based positioning algorithm at least includes: establishing a fingerprint database of all signal sources in the target area, and determining the location information of the terminal device by matching the wireless communication signal with the fingerprint database. Optionally, the location information can be two-dimensional location information or three-dimensional location information, which is not limited by the present invention.
[0109] Alternatively, the larger area where the target area is located can be pre-divided into multiple areas, and the target area and the area range corresponding to the target area can be determined from the multiple areas. For example, the larger area can be an electronic map generated by scanning the floor plan of the target family residence, and then divided into areas such as the master bedroom, bathroom, kitchen, living room, balcony, etc. according to living conditions. The living room where the OTT TV is installed can then be determined as the target area.
[0110] It can be seen that by implementing this optional implementation method, the location information of each terminal device can be determined based on the wireless communication signals of multiple terminal devices based on the signal positioning algorithm, and the number of signal-perceiving viewers in the target area can be determined based on whether the location information of each terminal device is located in the target area. This can accurately determine the number of viewers who perceive the signal dimension, which will help to accurately and comprehensively determine the number of viewers in the area in the subsequent analysis, and provide a real data basis for subsequent viewing indicator analysis.
[0111] As an optional implementation, in step 103, determining the number of viewers in the target area based on the number of visually perceived viewers and the number of signal-perceived viewers includes:
[0112] Determine whether the number of visually perceived viewers and the number of signal-perceived viewers are consistent;
[0113] When the judgment result is consistent, the number of visually perceived viewers or the number of signal-perceived viewers is determined as the number of viewers in the target area;
[0114] When the judgment result is inconsistent, the visually perceived number of viewers is corrected to obtain the number of viewers in the target area.
[0115] Optionally, the method for correcting the number of visually perceived viewers may be to calculate an average or weighted average of the number of visually perceived viewers and the number of signal perceived viewers to determine the number of viewers in the target area.
[0116] It can be seen that by implementing this optional implementation method, the method for calculating the number of viewers in the target area can be determined by judging whether the number of visually perceived viewers and the number of signal-perceived viewers are consistent, thereby being able to accurately and comprehensively determine the number of viewers in the area and provide a real data basis for subsequent viewing indicator analysis.
[0117] As an optional implementation, in the above step, correcting the number of visually perceived viewers to obtain the number of viewers in the target area includes:
[0118] determining whether the number of visually perceived viewers is greater than the number of signal-perceived viewers, and if so, determining the number of visually perceived viewers as the number of viewers in the target area;
[0119] If not, calculate the difference between the number of visually perceived viewers and the number of signal-perceived viewers;
[0120] Determine whether the difference is greater than a preset difference threshold;
[0121] If so, the number of visually perceived viewers is determined to be the number of viewers in the target area;
[0122] If not, the number of visually perceived viewers is corrected to obtain the number of viewers in the target area.
[0123] The purpose of the above setting is that the accuracy of the number of viewers perceived by visual perception is generally higher than the number of viewers perceived by signal. This is because the number of viewers informed by the signal dimension may include the number of viewers who are looking at the communication device but not watching, or the number of communication devices without owners. Therefore, when the number of visually perceived viewers is greater than the number of signal-perceived viewers, the number of visually perceived viewers shall prevail. However, when the number of visually perceived viewers is less than the number of signal-perceived viewers (the previous step has determined that the two cannot be consistent), it is necessary to calculate the difference between the two. When the difference between the two is large, it can be determined that the number of viewers who are looking at the communication device but not watching, or the number of communication devices without owners is higher. However, when the difference is small, there is reason to suspect that there may be a deviation in the number of visually perceived viewers, that is, the number of viewers calculated based on the audience image may have an algorithm-level error.
[0124] As an optional implementation, the visually perceived audience count is calculated by local devices deployed in the target area based on an offline algorithm model, while the signal-perceived audience count is calculated by cloud-based devices. In other words, in this implementation, audience image information is sent to the local device for visual perception to determine the audience count. This is generally based on an offline image recognition model, which has low accuracy, but effectively prevents audience images from being uploaded to the cloud, thus protecting user privacy. Device communication information, on the other hand, is sent to the cloud to calculate the audience count. This is because the privacy protection requirements for device communication information are relatively low, and confidentiality can be achieved through a certain level of low-cost encryption.
[0125] Accordingly, in the above step, correcting the number of visually perceived viewers to obtain the number of viewers in the target area may include:
[0126] Sending the audience image information to the cloud device so that the cloud device determines the number of cloud visually perceived audiences corresponding to the audience image information based on an online image recognition algorithm model;
[0127] Determine whether the number of cloud-based visually perceived viewers is consistent with the number of visually perceived viewers;
[0128] If the judgment result is no, the number of viewers in the target area is determined based on the number of cloud-based visual perception viewers and the number of signal perception viewers.
[0129] Specifically, the above steps can be regarded as a specific expansion of the step of "correcting the number of visually perceived viewers to obtain the number of viewers within the target area" in the above two embodiments, that is, this step can be the next step when it is judged that the number of visually perceived viewers and the number of signal-perceived viewers are inconsistent, or it can be the next step when it is judged that the difference between the number of visually perceived viewers and the number of signal-perceived viewers is less than or equal to a preset difference threshold.
[0130] Specifically, the purpose of the above steps is that when it is judged that the number of visually perceived viewers and the number of signal-perceived viewers are inconsistent, or when it is judged that the difference between the number of visually perceived viewers and the number of signal-perceived viewers is less than or equal to a preset difference threshold, it is generally tended to be believed that the number of viewers calculated by the local device is inaccurate. Therefore, the audience image information obtained locally can be uploaded to the cloud device to calculate the number of cloud-based visually perceived viewers. The cloud device generally has a higher recognition accuracy based on the online image recognition algorithm model, and can be used as a backup solution to prevent calculation errors.
[0131] Optionally, the number of viewers within the target area is determined based on the number of visually perceived viewers in the cloud and the number of signal-perceived viewers. The technical details of determining the number of viewers within the target area based on the number of visually perceived viewers and the number of signal-perceived viewers in the above embodiments can be referred to. That is, when the number of visually perceived viewers in the cloud is greater than or equal to the number of signal-perceived viewers, the number of visually perceived viewers in the cloud can be determined as the number of viewers within the target area, and when the number of visually perceived viewers in the cloud is less than the number of signal-perceived viewers, the number of visually perceived viewers in the cloud can be corrected to obtain the number of viewers within the target area. Similarly, the number of visually perceived viewers in the cloud can be corrected by calculating the average or weighted average of the number of visually perceived viewers in the cloud and the number of signal-perceived viewers to determine the number of viewers within the target area.
[0132] Example 2
[0133] See also Figure 2 , Figure 2 This is a schematic diagram of the structure of a device for determining the number of viewers based on hybrid perception disclosed in an embodiment of the present invention. Figure 2 The described apparatus for determining the number of viewers may be applied to a viewer number calculation chip, a calculation terminal or a calculation server (wherein the calculation server may be a local server or a cloud server). Figure 2 As shown, the apparatus for determining the number of viewers based on hybrid perception may include:
[0134] The visual perception module 201 is used to obtain audience image information of a target area, and determine the number of visually perceived audience members in the target area based on the audience image information.
[0135] Optionally, the target area may be an indoor area or an outdoor area. The target area may be equipped with a playback device for playing specific media, such as an OTT TV (Over The Top TV) or a large advertising screen. The method described in the present invention can then be used to analyze the number of viewers of the played specific media. For example, the target area may be a living room in a home equipped with a smart TV.
[0136] Alternatively, the audience image information may be acquired by an image acquisition device, such as a camera, located in the target area. Alternatively, the image acquisition device may be located on the aforementioned playback device, such as an OTTTV integrated with a camera.
[0137] Optionally, the number of visually perceived viewers in the target area may be determined based on the audience image information using an image recognition algorithm.
[0138] The signal sensing module 202 is configured to obtain device communication signals in a target area and determine the number of signal sensing viewers in the target area based on the device communication signals.
[0139] Optionally, the device communication signal may be a wired signal or a wireless signal. Optionally, the device communication signal may be a Bluetooth signal, a Wi-Fi signal, or other electromagnetic wave signal that can be used for communication. Optionally, the device communication signal may be received by a signal receiving device located in the target area and transmitted to a cloud server for analysis to determine the number of signal-aware viewers in the target area based on the device communication signal.
[0140] The number determination module 203 is configured to determine the number of viewers in the target area according to the number of visually perceived viewers and the number of signal-perceived viewers.
[0141] Optionally, the number of viewers in the target area may be determined by calculating the sum or difference between the number of visually perceived viewers and the number of signal-perceived viewers, or by calculating the number of visually perceived viewers and the number of signal-perceived viewers under specific rules.
[0142] Optionally, the number of audience members in the target area can be used to determine audience portrait information. For example, the portrait information of each audience member in the target area, such as age, occupation, hobbies, etc., can be determined as the audience portrait of the target area for subsequent data analysis.
[0143] It can be seen that the above-mentioned embodiments of the invention can use the number of viewers perceived by vision and the number of viewers perceived by signals to comprehensively determine the number of viewers in the area, thereby determining the number of viewers more accurately and efficiently, and providing a real data basis for subsequent viewing indicator analysis.
[0144] As an optional implementation, the visual perception module 201 determines the specific method of determining the number of visually perceived viewers in the target area based on the audience image information, including:
[0145] Determine facial features in the audience's image information based on image recognition algorithms;
[0146] Based on facial features, the number of visually perceived viewers in the target area is determined.
[0147] Optionally, the audience image information may be two-dimensional image information or three-dimensional image information, wherein the image recognition algorithm may be a two-dimensional face image recognition algorithm, such as a face template matching algorithm, an image matrix singular value feature recognition algorithm, a subspace analysis method, a local preserving projection algorithm, a principal component analysis method, an elastic matching method, a KL transform-based eigenface method, an artificial neural network method, a support vector machine method, an integral image feature method, or a probability model-based method. Optionally, the image recognition algorithm may also be a three-dimensional face image recognition algorithm, such as a 3D local matching method based on image features or a method based on model variable parameters, wherein the method based on model variable parameters combines the 3D deformation of a universal face model with matrix iterative minimization based on distance mapping to restore the head posture and 3D face, continuously updating the posture parameters as the correlation relationship of the model deformation changes, and repeating this process until the minimization scale meets the requirement.
[0148] Optionally, the number of visually perceived viewers in the target area may be determined based on the number of facial features in the target area.
[0149] It can be seen that by implementing this optional implementation method, the facial features in the audience image information can be determined based on the image recognition algorithm, and the number of visually perceived viewers in the target area can be determined based on the facial features, so that the number of viewers perceived in the visual dimension can be accurately determined, which is helpful for the subsequent accurate and comprehensive determination of the number of viewers in the area, and provides a real data basis for subsequent viewing indicator analysis.
[0150] As an optional implementation, the visual perception module 201 determines the number of visually perceived viewers in the target area according to facial features in a specific manner including:
[0151] For any facial feature, calculate the viewing intention of the facial feature;
[0152] Determining whether the viewing intention is within a preset viewing intention score range corresponding to the target area, and if the determination result is yes, determining the facial feature as a viewing facial feature;
[0153] The total number of all viewed facial features is determined as the number of visually perceived viewers in the target area.
[0154] Optionally, viewing intention may include one or a combination of facial orientation angle and eye gaze angle. For example, viewing intention may be the facial orientation angle detected by a facial orientation angle detection regression model, or the eye gaze angle detected by an eye gaze angle detection regression model, or the user viewing score detected by a user viewing detection regression model that can simultaneously detect facial orientation angle and eye gaze angle. Accordingly, the preset viewing intention score range may be a preset facial orientation angle range, eye gaze angle range, or user viewing score range.
[0155] Optionally, a face orientation angle detection regression model can be trained. As an example, ResNet18 can be used as the network model skeleton of the face orientation angle detection regression model. Collect and annotate face orientation angle data, use face images containing pre-annotated face orientation angle values (each face image includes face orientation angle values in three directions: pitch, yaw, and roll) as the first training sample, and use the first training sample to train the face orientation angle detection regression model to obtain a trained face orientation angle detection regression model. Specifically, during training, the first training sample is used as the input of the face posture regression model, and the output is the face orientation angle values in three directions: pitch, yaw, and roll. When the average error of the output angles in the three directions is less than the preset angle (such as 5 degrees), stop training.
[0156] Optionally, a human eye gaze angle detection regression model can be trained. As an example, VGG19 can be used as the network model skeleton of the human eye gaze angle detection regression model. Collect and annotate human eye gaze angle data, and use human eye images containing pre-annotated human eye gaze angle values (each eye image includes the human eye gaze angle values of the left eye in the three directions of pitch, yaw, and roll and the annotated values of the human eye gaze angle values of the right eye in the three directions of pitch, yaw, and roll) as the second training sample, and use the second training sample to train the human eye gaze angle detection regression model to obtain a trained human eye gaze angle detection regression model. Specifically, the second training sample is used as the input of the eye gaze angle detection regression model, and the output is a set of eye gaze angles in three directions (left eye + right eye pitch, yaw, and roll). That is, the output of the eye gaze angle detection regression model is six angle values in total, namely the eye gaze angle values in the three directions of pitch, yaw, and roll of the left eye and the eye gaze angle values in the three directions of pitch, yaw, and roll of the right eye. When the average error of the output angles in the three directions is less than 8 degrees, the training is stopped. It should be noted that the eye image containing the pre-marked eye gaze angles can be a whole face image collected in advance (that is, the eye gaze angle values are pre-marked on the face image), or only the eye area part in the face image can be input into the gaze regression model as the second training sample. Those skilled in the art can select the corresponding training sample according to actual needs.
[0157] Optionally, a user viewing detection regression model can be trained. Specifically, a collection device is first used to collect multiple face videos, each video is about 2 seconds (a total of about 60 frames of images). The collected data is labeled and divided into two categories: images of viewed faces and images of unviewed faces, and the labeled face images are used as the third training sample. Then, the third training sample is input into the above-mentioned trained face orientation angle detection regression model and the trained eye gaze angle detection regression model to obtain face orientation angle values (angle values in the three directions of pitch, yaw, and roll) and eye gaze angle values (left eye gaze angle values in the three directions of pitch, yaw, and roll and right eye gaze angle values in the three directions of pitch, yaw, and roll). In this way, 9 angle values can be obtained for each frame of face image. As an example, the angle values of 10 frames of images can be randomly extracted from each face video data as the input of the user viewing detection regression model. In this way, 90 angle values can be obtained for each face video data (10 frames of face images are randomly extracted from each video data, and 9 angle values are obtained for each frame of face image).
[0158] Optionally, a multi-layer perceptron (MLP) can be used to train a user viewing detection regression model for classification. The 90 angle values obtained from each face video data set are used as input features for the user viewing detection regression model. The model is trained and the output is the user viewing score. The trained model is tested. If it fails the test, the model is retrained until it passes the test, completing the training.
[0159] Specifically, when calculating the viewing willingness of facial features, the facial features can be input into a facial orientation angle detection regression model or a human eye gaze angle detection regression model to obtain corresponding angle values, and the corresponding angle values are used as the viewing willingness. Alternatively, multiple frames of facial images related to the facial features can be input into a trained facial orientation angle detection regression model, and each frame of facial image obtains facial orientation angle values in the three directions of pitch, yaw, and roll. At the same time, multiple frames of facial images can be input into a trained human eye gaze angle detection regression model, and each frame of facial image obtains left eye gaze angle values in the three directions of pitch, yaw, and roll, and right eye gaze angle values in the three directions of pitch, yaw, and roll. Then, the obtained facial orientation angle values and human eye gaze angle values of the multiple frames of facial images (a total of 90 angle values as an input feature) are input into a trained user viewing detection regression model to obtain a user viewing score, and the user viewing score is used as the viewing willingness of the facial feature.
[0160] As an optional embodiment, the device communication signal includes wireless communication signals of multiple terminal devices; the signal sensing module 202 determines the number of signal sensing viewers in the target area based on the device communication signal in a specific manner, including:
[0161] Based on the signal positioning algorithm, the location information of each terminal device is determined according to the wireless communication signals of multiple terminal devices;
[0162] For any terminal device, determine whether the location information is within the target area, and if the judgment result is yes, determine that the terminal device is a terminal within the area;
[0163] The total number of terminals in all areas is determined as the number of signal-perceiving viewers in the target area.
[0164] Optionally, the signal positioning algorithm may include at least one of a weight-based positioning algorithm, a fingerprint-based positioning algorithm, and a positioning method based on triangulation. Optionally, the calculation basis of the weight-based positioning algorithm includes at least one of a distance weight, a signal accuracy weight, and a signal source weight. Optionally, the positioning strategy of the fingerprint-based positioning algorithm at least includes: establishing a fingerprint database of all signal sources in the target area, and determining the location information of the terminal device by matching the wireless communication signal with the fingerprint database. Optionally, the location information can be two-dimensional location information or three-dimensional location information, which is not limited by the present invention.
[0165] Alternatively, the larger area where the target area is located can be pre-divided into multiple areas, and the target area and the area range corresponding to the target area can be determined from the multiple areas. For example, the larger area can be an electronic map generated by scanning the floor plan of the target family residence, and then divided into areas such as the master bedroom, bathroom, kitchen, living room, balcony, etc. according to living conditions. The living room where the OTT TV is installed can then be determined as the target area.
[0166] It can be seen that by implementing this optional implementation method, the location information of each terminal device can be determined based on the wireless communication signals of multiple terminal devices based on the signal positioning algorithm, and the number of signal-perceiving viewers in the target area can be determined based on whether the location information of each terminal device is located in the target area. This can accurately determine the number of viewers who perceive the signal dimension, which will help to accurately and comprehensively determine the number of viewers in the area in the subsequent analysis, and provide a real data basis for subsequent viewing indicator analysis.
[0167] As an optional implementation, the specific method of the number determination module 203 determining the number of viewers in the target area according to the number of visually perceived viewers and the number of signal-perceived viewers includes:
[0168] Determine whether the number of visually perceived viewers and the number of signal-perceived viewers are consistent;
[0169] When the judgment result is consistent, the number of visually perceived viewers or the number of signal-perceived viewers is determined as the number of viewers in the target area;
[0170] When the judgment result is inconsistent, the visually perceived number of viewers is corrected to obtain the number of viewers in the target area.
[0171] Optionally, the method for correcting the number of visually perceived viewers may be to calculate an average or weighted average of the number of visually perceived viewers and the number of signal perceived viewers to determine the number of viewers in the target area.
[0172] It can be seen that by implementing this optional implementation method, the method for calculating the number of viewers in the target area can be determined by judging whether the number of visually perceived viewers and the number of signal-perceived viewers are consistent, thereby being able to accurately and comprehensively determine the number of viewers in the area and provide a real data basis for subsequent viewing indicator analysis.
[0173] As an optional implementation, the specific method of the number determination module 203 correcting the visually perceived number of viewers to obtain the number of viewers in the target area includes:
[0174] determining whether the number of visually perceived viewers is greater than the number of signal-perceived viewers, and if so, determining the number of visually perceived viewers as the number of viewers in the target area;
[0175] If not, calculate the difference between the number of visually perceived viewers and the number of signal-perceived viewers;
[0176] Determine whether the difference is greater than a preset difference threshold;
[0177] If so, the number of visually perceived viewers is determined to be the number of viewers in the target area;
[0178] If not, the number of visually perceived viewers is corrected to obtain the number of viewers in the target area.
[0179] The purpose of the above setting is that the accuracy of the number of viewers perceived by visual perception is generally higher than the number of viewers perceived by signal. This is because the number of viewers informed by the signal dimension may include the number of viewers who are looking at the communication device but not watching, or the number of communication devices without owners. Therefore, when the number of visually perceived viewers is greater than the number of signal-perceived viewers, the number of visually perceived viewers shall prevail. However, when the number of visually perceived viewers is less than the number of signal-perceived viewers (the previous step has determined that the two cannot be consistent), it is necessary to calculate the difference between the two. When the difference between the two is large, it can be determined that the number of viewers who are looking at the communication device but not watching, or the number of communication devices without owners is higher. However, when the difference is small, there is reason to suspect that there may be a deviation in the number of visually perceived viewers, that is, the number of viewers calculated based on the audience image may have an algorithm-level error.
[0180] As an optional implementation, the visually perceived audience number is calculated by the visual perception module 201 triggering a local device deployed in the target area based on an offline algorithm model, while the signal-perceived audience number is calculated by the signal perception module 202 triggering a cloud device. That is, in this implementation, the audience image information is sent to the local device for visual perception to determine the audience number. This is generally performed based on an offline image recognition model, which has low accuracy, but can effectively ensure that the audience image is not uploaded to the cloud to protect user privacy. Device communication information is sent to the cloud to calculate the audience number. This is because the privacy protection requirements for device communication information are not high and can be achieved through a certain degree of low-cost encryption.
[0181] Accordingly, the specific method of the number determination module 203 correcting the visually perceived number of viewers to obtain the number of viewers in the target area includes:
[0182] Sending the audience image information to the cloud device so that the cloud device determines the number of cloud visually perceived audiences corresponding to the audience image information based on an online image recognition algorithm model;
[0183] Determine whether the number of cloud-based visually perceived viewers is consistent with the number of visually perceived viewers;
[0184] If the judgment result is no, the number of viewers in the target area is determined based on the number of cloud-based visual perception viewers and the number of signal perception viewers.
[0185] Specifically, the above specific method can be regarded as a specific expansion of the specific method of "correcting the number of visually perceived viewers to obtain the number of viewers in the target area" in the above two implementations, that is, this specific method can be the next step when the number determination module 203 determines that the number of visually perceived viewers and the number of signal-perceived viewers are inconsistent, or it can be the next step when the number determination module 203 determines that the difference between the number of visually perceived viewers and the number of signal-perceived viewers is less than or equal to the preset difference threshold.
[0186] Specifically, the purpose of the above-mentioned specific method is that when it is judged that the number of visually perceived viewers and the number of signal-perceived viewers are inconsistent, or when it is judged that the difference between the number of visually perceived viewers and the number of signal-perceived viewers is less than or equal to a preset difference threshold, it is generally tended to be believed that the number of viewers calculated by the local device is inaccurate. Therefore, the audience image information obtained locally can be uploaded to the cloud device to calculate the number of cloud-based visually perceived viewers. The cloud device generally has a higher recognition accuracy based on the online image recognition algorithm model, and can be used as a backup solution to prevent calculation errors.
[0187] Optionally, the number of viewers within the target area is determined based on the number of visually perceived viewers in the cloud and the number of signal-perceived viewers. The technical details of determining the number of viewers within the target area based on the number of visually perceived viewers and the number of signal-perceived viewers in the above embodiments can be referred to. That is, when the number of visually perceived viewers in the cloud is greater than or equal to the number of signal-perceived viewers, the number of visually perceived viewers in the cloud can be determined as the number of viewers within the target area, and when the number of visually perceived viewers in the cloud is less than the number of signal-perceived viewers, the number of visually perceived viewers in the cloud can be corrected to obtain the number of viewers within the target area. Similarly, the number of visually perceived viewers in the cloud can be corrected by calculating the average or weighted average of the number of visually perceived viewers in the cloud and the number of signal-perceived viewers to determine the number of viewers within the target area.
[0188] Example 3
[0189] See also Figure 3 , Figure 3 This is another apparatus for determining the number of viewers based on hybrid perception disclosed in an embodiment of the present invention. Figure 3 The described apparatus for determining the number of viewers may be applied to a viewer number calculation chip, a calculation terminal or a calculation server (wherein the calculation server may be a local server or a cloud server). Figure 3 As shown, the apparatus for determining the number of viewers based on hybrid perception may include:
[0190] A memory 301 storing executable program code;
[0191] a processor 302 coupled to the memory 301;
[0192] The processor 302 calls the executable program code stored in the memory 301 to execute the steps of the method for determining the number of viewers based on hybrid perception described in the first or second embodiment.
[0193] Example 4
[0194] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the method for determining the number of viewers based on hybrid perception described in the first or second embodiment.
[0195] Example 5
[0196] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps of the method for determining the number of viewers based on hybrid perception described in Example 1 or Example 2.
[0197] The foregoing description of specific embodiments of the present disclosure is intended to illustrate a method for performing a multi-tasking process. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0198] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer-readable storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant details, refer to the descriptions of the method embodiments.
[0199] The apparatus, device, non-volatile computer-readable storage medium and method provided in the embodiments of this specification correspond to each other. Therefore, the apparatus, device, and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device, and non-volatile computer storage medium will not be repeated here.
[0200] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using a hardware module. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, eliminating the need for a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0201] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0202] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0203] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0204] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0205] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0206] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0207] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0208] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0209] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0210] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0211] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0212] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0213] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0214] Finally, it should be noted that the method and device for determining the number of viewers based on hybrid perception disclosed in the embodiments of the present invention are only preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for determining the number of viewers based on hybrid perception, characterized in that: The method comprises: Acquiring audience image information of a target area, and determining the number of visually perceived audience members in the target area based on the audience image information; Acquiring device communication signals in the target area, and determining the number of signal-perceiving viewers in the target area based on the device communication signals; The number of viewers in the target area is determined according to the number of visually perceived viewers and the number of signal-perceived viewers.
2. The method for determining the number of viewers based on hybrid perception according to claim 1, characterized in that: Determining the number of visually perceived viewers in the target area based on the audience image information includes: Determining facial features in the audience image information based on an image recognition algorithm; The number of visually perceived viewers in the target area is determined based on the facial features.
3. The method for determining the number of viewers based on hybrid perception according to claim 2, characterized in that: Determining the number of visually perceived viewers within the target area based on the facial features includes: For any of the facial features, calculating the viewing intention of the facial feature; determining whether the viewing intention is within a preset viewing intention score range corresponding to the target area, and if the judgment result is yes, determining the facial feature as a viewing facial feature; The total number of all viewed facial features is determined as the number of visually perceived viewers in the target area.
4. The method for determining the number of viewers based on hybrid perception according to claim 1, characterized in that: The device communication signal includes wireless communication signals of multiple terminal devices; and determining the number of signal-perceiving viewers in the target area based on the device communication signal includes: Determine the location information of each terminal device based on the wireless communication signals of the plurality of terminal devices based on a signal positioning algorithm; For any of the terminal devices, determining whether the location information is within the target area, and if the determination result is yes, determining that the terminal device is a terminal within the area; The total number of terminals in all the areas is determined as the number of signal-perceiving viewers in the target area.
5. The method for determining the number of viewers based on hybrid perception according to claim 1, characterized in that: The determining the number of viewers in the target area according to the number of visually perceived viewers and the number of signal-perceived viewers includes: Determining whether the number of visually perceived viewers is consistent with the number of signal-perceived viewers; When the judgment result is consistent, determining the number of visually perceived viewers or the number of signal-perceived viewers as the number of viewers in the target area; When the judgment result is inconsistent, the visually perceived number of viewers is corrected to obtain the number of viewers in the target area.
6. The method for determining the number of viewers based on hybrid perception according to claim 5, characterized in that: The correcting the visually perceived number of viewers to obtain the number of viewers in the target area includes: Determining whether the number of visually perceived viewers is greater than the number of signal-perceived viewers, and if so, determining the number of visually perceived viewers as the number of viewers in the target area; If not, calculating the difference between the number of visually perceived viewers and the number of signal-perceived viewers; Determining whether the difference is greater than a preset difference threshold; If yes, determining the number of visually perceived viewers as the number of viewers in the target area; If not, the visually perceived number of viewers is corrected to obtain the number of viewers in the target area.
7. The method for determining the number of viewers based on hybrid perception according to claim 5 or 6, characterized in that: The number of visually perceived viewers is calculated by local devices deployed in the target area based on an offline algorithm model; the number of signal-perceived viewers is calculated by cloud devices; The correcting the visually perceived number of viewers to obtain the number of viewers in the target area includes: Sending the audience image information to the cloud device, so that the cloud device determines the number of cloud visually perceived audiences corresponding to the audience image information based on an online image recognition algorithm model; Determining whether the number of cloud-based visually perceived viewers is consistent with the number of visually perceived viewers; If the judgment result is no, the number of viewers in the target area is determined according to the number of cloud-based visual perception viewers and the number of signal perception viewers.
8. A device for determining the number of viewers based on hybrid perception, characterized in that: The device comprises: A visual perception module, configured to obtain image information of audience members in a target area, and determine the number of visually perceived audience members in the target area based on the image information of the audience members; a signal sensing module, configured to obtain device communication signals in the target area, and determine the number of signal sensing viewers in the target area based on the device communication signals; The number determination module is used to determine the number of audience members in the target area according to the number of visually perceived audience members and the number of signal-perceived audience members.
9. A device for determining the number of viewers based on hybrid perception, characterized in that: The device comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the method for determining the number of viewers based on hybrid perception according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer program for electronic data exchange is stored therein, wherein the computer program enables a computer to execute the method for determining the number of audience members based on mixed perception according to any one of claims 1 to 7.
Citation Information
Patent Citations
Naked eye 3D display method and apparatus
CN107167926A
Television viewing user counting method and counting system and intelligent television viewing terminal
CN108307233A