Facial expression material collection method and system
By using a circular arrangement of cameras to play video clips and automatically recording facial expression data based on eye gaze and facial expression changes, the problem of obtaining a facial expression database is solved, thus improving the accuracy of facial expression recognition and the accuracy of emotion recognition models in the intelligent cockpit.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZEBRED NETWORK TECH CO LTD
- Filing Date
- 2023-08-10
- Publication Date
- 2026-07-03
AI Technical Summary
In existing technologies, it is difficult to obtain a database of facial expressions under natural and realistic conditions, resulting in limited accuracy of facial expression recognition in smart cockpits.
By controlling several cameras arranged in a ring around the shooting point, video clips are played and facial expressions of the target object are captured. Using eye gaze direction and facial expression changes as trigger conditions, facial expression materials are automatically recorded. Combined with emotion tags and mapping relationships, comprehensive and automated facial expression material collection is achieved.
It improves the accuracy of facial expression recognition in intelligent cockpits, collects real and effective facial expression data, reduces the subjectivity of human annotation, and enhances the accuracy of emotion recognition models.
Smart Images

Figure CN116958502B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of intelligent cockpit technology, and in particular to a method and system for collecting facial expression data. Background Technology
[0002] Currently, deep learning-based emotion recognition technology is developing rapidly and has made significant progress in the application of smart cockpits. Since facial expression emotion recognition is the most practical emotion recognition technology, the future research direction for smart cockpits is to enable AI (artificial intelligence) robots to overcome the limitations of object recognition and increase their emotion perception capabilities.
[0003] In the realm of deep learning, facial expression-based emotion recognition technology generally involves two key steps: building a large-scale facial expression database and training an emotion recognition network. Because obtaining a database of natural, real-world facial expressions is extremely difficult, it limits the accuracy of related emotion recognition models.
[0004] Therefore, how to automatically collect real and effective facial expression data to improve the accuracy of facial expression recognition for vehicle users in smart cockpits is an urgent problem to be solved. Summary of the Invention
[0005] To address the aforementioned technical problems, the first aspect of this specification discloses a method for acquiring facial expression data, the method comprising:
[0006] Control several camera devices to be arranged in a ring around the shooting point, and control the several playback devices to play video clips;
[0007] When the target object at the shooting location watches the video clip, the target object's facial expressions are captured and recognized.
[0008] The first trigger condition is a change in the facial expression of the target object, which triggers the target camera device among the plurality of camera devices to start recording the facial expression material of the target object.
[0009] Preferably, the plurality of camera devices have their own device numbers, the plurality of camera devices and the plurality of playback devices are paired, and the paired camera devices and playback devices are bound to the same location.
[0010] Preferably, the video clip is configured with emotion tags for the video content; before controlling the plurality of playback devices to play the video clip, the method further includes:
[0011] A mapping relationship between video clips and device numbers is established to facilitate the lookup of the emotion tags configured for the video clips using the device numbers.
[0012] Preferably, when a target object at the shooting location views the video clip, the method further includes:
[0013] Track the direction of the target object's eye gaze, identify the camera device located in the direction of the target object's eye gaze as the target camera device, and obtain the target number of the target camera device.
[0014] Preferably, after obtaining the target number of the target camera device, the method further includes:
[0015] Based on the target number and the mapping relationship, the emotion tag corresponding to the facial expression material is determined.
[0016] Preferably, after triggering the target camera among the plurality of camera devices to start recording the facial expression material of the target object, the method further includes:
[0017] Continuously track the direction of the target object's eye gaze;
[0018] The second trigger condition is that the target object's eye gaze leaves the target camera device, triggering the target camera device to stop recording the facial expression material.
[0019] Preferably, after triggering the target camera among the plurality of camera devices to start recording the facial expression material of the target object, the method further includes:
[0020] Continuously analyze the facial expressions of the target object;
[0021] The third trigger condition is that the target camera device is simultaneously triggered to stop recording the facial expression material when the target object's facial expression changes again, and then restarts recording the facial expression material after the target object's facial expression changes again.
[0022] Preferably, after triggering the target camera among the plurality of camera devices to start recording the facial expression material of the target object, the method further includes:
[0023] Control all camera devices other than the target camera device to simultaneously start recording the facial expression material; and simultaneously stop recording the facial expression material when the target camera device finishes recording the facial expression material;
[0024] The emotion tags are synchronously configured in the facial expression footage recorded by the other camera devices.
[0025] A second aspect of this specification discloses a facial expression data acquisition system, comprising: a plurality of camera devices, a plurality of playback devices, and a control device; wherein the plurality of camera devices are arranged in a ring around the shooting point;
[0026] The control device specifically includes:
[0027] The control module is used to control the plurality of playback devices to play video clips;
[0028] The recognition module is used to capture and recognize the facial expressions of the target object when the target object at the shooting location watches the video clip;
[0029] The recording module is used to trigger the target camera device among the plurality of camera devices to start recording the facial expression material of the target object when the target object's facial expression changes as the first trigger condition.
[0030] Preferably, the plurality of camera devices have their own device numbers, the plurality of camera devices and the plurality of playback devices are paired, and the paired camera devices and playback devices are bound to the same location.
[0031] Preferably, the control device further includes:
[0032] A construction module is used to build a mapping relationship between video segments and device numbers, so as to use the device number to find the emotion tag configured for the video segment.
[0033] Preferably, the control device further includes:
[0034] The eye-tracking module is used to track the direction of the target object's eye gaze, identify the camera device in the direction of the target object's eye gaze as the target camera device, and obtain the target number of the target camera device.
[0035] Preferably, the control device further includes:
[0036] The determination module is used to determine the emotion tag corresponding to the facial expression material based on the target number and the mapping relationship.
[0037] Preferably, the gaze tracking module is further configured to continuously track the gaze direction of the target object;
[0038] The recording module is further configured to trigger the target camera device to end the recording of the facial expression material when the target object's eye gaze leaves the target camera device as a second trigger condition.
[0039] Preferably, the recognition module is further configured to continuously analyze the facial expressions of the target object;
[0040] The recording module is also used to trigger the target camera device to stop recording the facial expression material and restart recording the facial expression material after the target object's facial expression changes again, using the third trigger condition of the target object's facial expression changing again.
[0041] Preferably, the control device further includes:
[0042] The control module is also used to control other camera devices besides the target camera device to simultaneously start recording the facial expression material; and to simultaneously stop recording the facial expression material when the target camera device finishes recording the facial expression material;
[0043] The configuration module is used to synchronously configure the emotion tags in the facial expression materials recorded by the other camera devices.
[0044] A third aspect of this specification discloses a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the above-described method.
[0045] A fourth aspect of this specification discloses a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method.
[0046] Through one or more embodiments of this specification, this specification has the following beneficial effects or advantages:
[0047] The technical solution in this specification constructs a realistic scene for capturing facial expression data from multiple angles by controlling several camera devices arranged in a ring around a shooting point and controlling several playback devices to play video clips. This induces a genuine emotional response from the target object at the shooting point when watching the video clips. Based on this, facial expression recognition is performed on the target object, capturing facial expression data that expresses the target object's true emotions when their facial expression changes due to the video content. This helps improve the accuracy of facial expression recognition for vehicle users in intelligent cockpits.
[0048] The above description is merely an overview of the technical solution in this specification. In order to better understand the technical means in this specification and to implement it in accordance with the contents of this specification, and to make the above and other objects, features and advantages of this specification more apparent and understandable, specific embodiments of this specification are given below. Attached Figure Description
[0049] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit this specification. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0050] Figure 1 A flowchart illustrating a method for acquiring facial expression data according to one embodiment of this specification is shown;
[0051] Figure 2 A schematic diagram of a real-world scene for capturing facial expression data from multiple angles, according to one embodiment of this specification, is shown.
[0052] Figure 3 A schematic diagram of a facial expression data acquisition system according to one embodiment of this specification is shown. Detailed Implementation
[0053] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0054] This specification provides a method for acquiring facial expression data. This method is applicable to acquiring human facial expression data, and also to acquiring facial expression data of animals with rich emotional expression. See also... Figure 1 The method includes the following steps:
[0055] Step 101: Control several camera devices to be arranged in a ring around the shooting point, and control several playback devices to play video clips.
[0056] In order to collect authentic and effective facial expression data, this solution pre-constructs a realistic collection scenario that facilitates the collection of facial expression data from multiple angles.
[0057] In this data collection scenario, several camera devices 201 and several playback devices 202 are configured, see [link / reference] Figure 2Several camera devices 201 are arranged in a ring around the shooting point. When the target object is at the shooting point, the ring-arranged camera devices 201 can capture the facial expressions of the target object from various angles, achieving all-round coverage of the target object's facial expressions. In this embodiment, the several camera devices 201 and several playback devices 202 are paired. The playback devices 202 are used to play video clips, such as funny videos, sad videos, etc. It is worth noting that each video clip is assigned an emotional tag. For example, funny video clips are assigned emotional tags such as funny, happy, excited, etc., while sad videos are assigned emotional tags such as crying, sad, depressed, sad, etc. The video clips and playback devices are played in a random combination. For example, five playback devices and five video clips with different emotional tags can be used. For instance, the same playback device can play five video clips sequentially; five playback devices can play the same video clip simultaneously; or five playback devices can play video clips with different emotional tags simultaneously; or only one playback device can be selected to play five video clips sequentially, etc. The main purpose of playing video clips in this embodiment is to induce the target object to produce realistic and effective facial expressions, therefore the method of playing video clips is not limited.
[0058] The camera device 201 is used to capture the realistic and effective facial expressions of the target subject while watching the video. To more clearly capture the genuine emotions displayed by the target subject while watching the video clip, the paired camera device 201 and playback device 202 are attached to the same position, enabling frontal capture of the target subject's facial expressions. In the pairing relationship, several camera devices 201 are paired one-to-one with several playback devices 202; alternatively, there can be a many-to-few pairing relationship, such as five camera devices paired with three playback devices, where three camera devices are paired one-to-one with three playback devices, and the other two camera devices are arranged separately. Each camera device has its own device number, such as 001, 002, and so on.
[0059] In one alternative implementation, existing technologies for annotating facial expression images mostly rely on manual annotation, which is subjective and limits the accuracy of research or deep learning training of facial expression recognition methods. Therefore, in this embodiment, each video segment is equipped with an emotion tag specific to the video content. This emotion tag is obtained by analyzing the video content using relevant algorithms and is an accurate representation of the video content. A mapping relationship between video segments and device numbers is pre-established to facilitate finding the emotion tag configured for the video segment using the device number. Since the recorded facial expression material represents the target audience's genuine emotional response after watching the video segment content, the emotion tags configured for the video segment can be directly used as the emotion tags for the facial expression material, eliminating the need for further manual annotation and avoiding the subjectivity of manual annotation.
[0060] Step 102: When the target object at the shooting location is watching the video clip, capture and recognize the target object's facial expressions.
[0061] To accurately capture the target audience's genuine emotional responses while watching video clips, two methods are employed. First, the target audience's eye gaze direction is tracked, and the camera device positioned in that direction is designated as the target camera device. This target camera device is used to capture the target audience's genuine emotional responses. Upon identifying the target camera device, its target number is simultaneously obtained for subsequent use. Second, the target audience's facial expressions need to be captured and recognized to determine the triggering conditions for the target camera device to record facial expression footage.
[0062] This embodiment provides two methods for determining the target camera device. In one optional implementation, images containing head or eye features or head posture features are captured by several camera devices. Computer vision or deep learning techniques are used to extract and analyze these head or eye features or head posture features from the images to obtain approximate values for the eye gaze direction or head rotation angle. Combining this with the circular layout of the camera devices, the target camera device indicated by the eye gaze direction and its target number can be determined. In another optional implementation, eye-tracking technology is used to track the video stream captured by several camera devices targeting the target object, and the eye gaze direction of the target object is calculated. Combining this with the circular layout of the camera devices, the target camera device indicated by the eye gaze direction and its target number can be determined.
[0063] This embodiment uses facial expression recognition technology combined with a target camera device to acquire facial expression images in real time for recognition. Specifically, computer vision or deep learning technology is used to perform face detection and key point localization on the real-time acquired facial expression images. These key points typically include facial features such as eyes, mouth, and eyebrows. In each frame, the positions of facial key points are extracted and tracked, and changes in expression are identified and analyzed based on changes in key point positions. Using a pre-trained facial expression recognition model or a custom model, the positions of facial key points and expressions are correlated and analyzed by calculating the motion trajectory, shape changes, and color or texture changes in specific areas of the key points. For continuous expression changes, more refined and accurate expression information can be obtained through analysis of consecutive frame images, and different time windows or filtering methods can be used to capture the trend and details of expression changes. When a change in facial expression is detected, such as a specific expression or an expression intensity reaching a certain threshold, the target camera device is triggered to execute step 103.
[0064] Step 103: Using a change in the facial expression of the target object as the first trigger condition, the target camera device among several camera devices is triggered to start recording the facial expression material of the target object.
[0065] In this embodiment, using changes in facial expression as the first trigger condition can avoid recording useless neutral expressions, making the recorded facial expression material realistic and effective.
[0066] During the recording process, the target camera device can be directly triggered to start recording facial expression footage. Alternatively, a shutter button can be installed in the target camera device, which can be automatically triggered or manually triggered to start recording facial expression footage based on a first trigger condition.
[0067] To determine the end of recording, the direction of the target subject's eye gaze is continuously tracked. A second trigger condition is used: when the target subject's eye gaze leaves the target camera device, the target camera device is triggered to stop recording facial expression footage. Specifically, when the target subject's eye gaze shifts from the currently recording target camera device to another camera device, the target camera device automatically ends recording. In this solution, the shift in eye gaze direction is used as the trigger condition to end the recording of facial expression footage, achieving fully automated and seamless recording.
[0068] The aforementioned method applies to scenarios where the target's gaze shifts. For scenarios where the target's gaze does not shift, the recording should be terminated as follows.
[0069] In one optional implementation, the end time of the video clip's playback is used as the recording end time for the target camera device, triggering the target camera device to stop recording facial expression material. For example, a comedy video clip has a playback duration of 5 minutes. If the target subject watches the entire video clip without any expression before watching, the recording starts when the target subject is detected laughing. If the target subject's gaze remains fixed throughout the clip, the recording ends when the comedy video clip's playback ends, thus recording facial expression material.
[0070] In one optional implementation, the facial expressions of the target subject are continuously analyzed; the specific analysis method has been described above and will not be repeated here. A third trigger condition is used to trigger the target camera to simultaneously stop recording facial expression material and restart recording the facial expression material after the target subject's facial expression changes again. Specifically, if video clips with different emotional tags are played consecutively on the same playback device, different emotional changes will be triggered in the target subject. Therefore, when the target subject's gaze does not shift, the change in the target subject's facial expression can be used as both the start and end trigger conditions for recording facial expression material. Recording of facial expression material begins when the target subject's facial expression changes, and the target subject's facial expression is continuously analyzed. After recording for a period of time, if the analysis shows that the target subject's facial expression is different from the facial expression captured at the previous moment (i.e., it has changed), the target camera is triggered to stop recording. Thus, this solution records facial expression material based on changes in the target subject's facial expression, and different emotional tags will be recorded when watching different types of video clips. Furthermore, multiple facial expression clips may be recorded from the same video clip, but these facial expression clips will have the same emotional tag.
[0071] Based on the obtained facial expression materials, the corresponding emotional tags are found in the mapping relationship using the target number, which can be used as the emotional tags corresponding to the facial expression materials, thereby avoiding the need for manual labeling of the facial expression materials.
[0072] The aforementioned scheme primarily utilizes a target camera device to record facial expression footage. To expand the data dimensions of facial expression footage and improve the training accuracy of facial expression recognition, in some optional implementations, other camera devices besides the target camera device are controlled to simultaneously begin recording facial expression footage; and when the target camera device finishes recording facial expression footage, the recording of facial expression footage also ends simultaneously; then, emotion tags are synchronously configured into the facial expression footage recorded by the other camera devices. In this scheme, using the start and end times of recording by the target camera device as trigger conditions, other camera devices are linked to capture facial expression footage of the user's genuine emotional responses from different angles, which expands the data dimensions of facial expression footage and lays a solid foundation for subsequently improving the training accuracy of facial expression recognition.
[0073] In this specification, real and effective facial expression materials are collected using the aforementioned disclosed method for training, resulting in an emotion recognition model that meets the accuracy requirements, which can then be deployed in an in-vehicle smart cockpit.
[0074] Based on the same inventive concept as the foregoing embodiments, the following implementation discloses a facial expression material acquisition system, including: a plurality of camera devices 201, a plurality of playback devices 202, and a control device 203; wherein, the plurality of camera devices 201 are arranged in a ring around a shooting point. The playback devices are used to play video clips, such as funny videos, sad videos, etc. The camera devices are used to capture real and valid facial expressions of the target object while watching the video. Optionally, the plurality of camera devices 201 are connected to the control device 203 via the Internet using the plurality of playback devices 202, and the control device 203 controls the playback of video clips by the plurality of playback devices 202 and the acquisition of facial expression materials by the plurality of camera devices 201. Optionally, the plurality of camera devices 201 are connected via the Internet using the plurality of playback devices 202. The plurality of playback devices 202 play video clips according to a pre-built playback process. The control device 203 acts as a control chip inside the plurality of camera devices 201, controlling the acquisition of facial expression materials by each camera device 201.
[0075] The control device 203 specifically includes:
[0076] Control module 301 is used to control the plurality of playback devices 202 to play video clips;
[0077] The recognition module 302 is used to capture and recognize the facial expressions of the target object when the target object at the shooting point watches the video clip;
[0078] The recording module 303 is used to trigger the target camera device among the plurality of camera devices 201 to start recording the facial expression material of the target object when the facial expression of the target object changes as the first trigger condition.
[0079] In some alternative implementations, a plurality of camera devices 201 have their own device numbers, the plurality of camera devices 201 and the plurality of playback devices 202 are paired, and the paired camera devices and playback devices are bound to the same location.
[0080] In some alternative embodiments, the control device 203 further includes:
[0081] A construction module is used to build a mapping relationship between video segments and device numbers, so as to use the device number to find the emotion tag configured for the video segment.
[0082] In some alternative embodiments, the control device 203 further includes:
[0083] The eye-tracking module is used to track the direction of the target object's eye gaze, identify the camera device in the direction of the target object's eye gaze as the target camera device, and obtain the target number of the target camera device.
[0084] In some alternative embodiments, the control device 203 further includes:
[0085] The determination module is used to determine the emotion tag corresponding to the facial expression material based on the target number and the mapping relationship.
[0086] In some alternative implementations,
[0087] The gaze tracking module is also used to continuously track the direction of the target object's eye gaze;
[0088] The recording module 303 is further configured to trigger the target camera device to end the recording of the facial expression material when the target object's eye gaze direction leaves the target camera device as a second trigger condition.
[0089] In some alternative implementations,
[0090] The recognition module 302 is also used to continuously analyze the facial expressions of the target object;
[0091] The recording module 303 is also used to trigger the target camera device to stop recording the facial expression material and restart recording the facial expression material after the target object's facial expression changes again, using the third trigger condition of the target object's facial expression changing again.
[0092] In some alternative embodiments, the control device 203 further includes:
[0093] The control module 301 is also used to control other camera devices besides the target camera device to simultaneously start recording the facial expression material; and to simultaneously stop recording the facial expression material when the target camera device finishes recording the facial expression material;
[0094] The configuration module is used to synchronously configure the emotion tags in the facial expression materials recorded by the other camera devices.
[0095] Based on the same inventive concept as the foregoing embodiments, the following implementation discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps disclosed in any of the foregoing method embodiments.
[0096] Based on the same inventive concept as the foregoing embodiments, the following implementation discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps disclosed in any of the foregoing method embodiments.
[0097] Through one or more embodiments of this specification, this specification has the following beneficial effects or advantages:
[0098] The technical solution in this specification constructs a realistic scene for capturing facial expression data from multiple angles by controlling several camera devices 201 arranged in a ring around a shooting point and controlling several playback devices 202 to play video clips. This induces a genuine emotional response from the target object at the shooting point when watching the video clips. Based on this, facial expression recognition is performed on the target object. When the target object's facial expression changes due to the video content, facial expression data that expresses the target object's true emotions is captured, thereby helping to improve the accuracy of facial expression recognition for vehicle users in the intelligent cockpit.
[0099] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, this specification is not directed to any particular programming language. It should be understood that the contents of this specification can be implemented using various programming languages, and the above descriptions of specific languages are for the purpose of disclosing preferred embodiments of this specification.
[0100] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this specification may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0101] Similarly, it should be understood that, in order to streamline this disclosure and aid in understanding one or more of the various inventive aspects, in the foregoing description of exemplary embodiments of this specification, various features of this specification are sometimes grouped together in a single embodiment, figure, or description thereof. However, this approach to disclosure should not be construed as reflecting an intention that the claimed specification requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this specification.
[0102] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0103] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this specification and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0104] The various component embodiments of this specification can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components of the gateway, proxy server, or system according to embodiments of this specification. This specification can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing some or all of the methods described herein. Such implementations of this specification can be stored on a computer-readable medium or can take the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0105] It should be noted that the above embodiments are illustrative of this specification and not limiting of it, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This specification can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
Claims
1. A method for collecting facial expression data, the method comprising: Several camera devices are controlled to be arranged in a ring around a shooting point, and several playback devices are controlled to play video clips. Each camera device has its own device number, and the camera devices and playback devices are paired, with the paired devices bound to the same location. A mapping relationship between video clips and device numbers is pre-established to facilitate finding the sentiment tags configured for the video clips using the device numbers. The video clips are configured with sentiment tags specific to the video content, and these sentiment tags are obtained by analyzing the video content using relevant algorithms. When the target object at the shooting location watches the video clip, the facial expression of the target object is captured and identified to obtain the trigger condition for the target camera device to record facial expression material; the direction of the target object's eye gaze is tracked, the camera device in the direction of the target object's eye gaze is identified as the target camera device, and the target number of the target camera device is obtained; The first trigger condition is a change in the facial expression of the target object, which triggers the target camera device among the plurality of camera devices to start recording the facial expression material of the target object; Based on the target number and the mapping relationship, the emotion tag corresponding to the facial expression material is determined; wherein, the emotion tag of the facial expression material is the emotion tag configured for the video segment.
2. The method of claim 1, wherein after triggering the target camera device among the plurality of camera devices to start recording the facial expression material of the target object, the method further comprises: Continuously track the direction of the target object's eye gaze; The second trigger condition is that the target object's eye gaze leaves the target camera device, triggering the target camera device to stop recording the facial expression material.
3. The method of claim 1, wherein after triggering the target camera device among the plurality of camera devices to start recording the facial expression material of the target object, the method further comprises: Continuously analyze the facial expressions of the target object; The third trigger condition is that the target camera device is simultaneously triggered to stop recording the facial expression material when the target object's facial expression changes again, and then restarts recording the facial expression material after the target object's facial expression changes again.
4. The method of claim 3, wherein after triggering the target camera among the plurality of camera devices to start recording the facial expression material of the target object, the method further comprises: Control all camera devices other than the target camera device to simultaneously begin recording the facial expression footage; And when the target camera device finishes recording the facial expression material, the recording of the facial expression material shall be stopped simultaneously; The emotion tags are synchronously configured in the facial expression footage recorded by the other camera devices.
5. A facial expression data acquisition system, comprising: A plurality of camera devices, a plurality of playback devices, and a control device; wherein, the plurality of camera devices are arranged in a ring around the shooting point; the plurality of camera devices have their own device numbers, the plurality of camera devices and the plurality of playback devices are paired, and the paired camera devices and playback devices are bound to the same position; The control device specifically includes: A control module is used to control the plurality of playback devices to play video segments; wherein, a mapping relationship between video segments and device numbers is pre-built so as to use the device number to find the sentiment tags configured for the video segment, the video segment is configured with sentiment tags for video content, and the sentiment tags are obtained by analyzing the video content using relevant algorithms; The recognition module is used to capture and recognize the facial expressions of the target object when the target object at the shooting point watches the video clip, so as to obtain the triggering condition for the target camera device to record facial expression material; track the direction of the target object's eye gaze, identify the camera device where the target object's eye gaze direction is located as the target camera device, and obtain the target number of the target camera device; The recording module is used to trigger the target camera device among the plurality of camera devices to start recording the facial expression material of the target object when the facial expression of the target object changes as the first trigger condition; The determination module is used to determine the emotion tag corresponding to the facial expression material based on the target number and the mapping relationship; wherein the emotion tag of the facial expression material is the emotion tag configured for the video segment.
6. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method according to any one of claims 1-4.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
Expression image marking method and system
CN106341724A
Facial image acquisition method, device, system and equipment and storage medium
CN116206353A