Method and system for face modeling

CN117173762BActive Publication Date: 2026-09-25SNAP INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311030981.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-10-19
Filing Date
2017-10-19
Publication Date
2026-09-25
Estimated Expiration
2037-10-19

Smart Images

  • Figure CN117173762B_ABST
    Figure CN117173762B_ABST
Patent Text Reader

Abstract

Systems, apparatuses, media, and methods are provided for modeling facial representations using image segmentation with client devices. The systems and methods receive an image depicting a face, detect at least a portion of the face within the image, and identify a set of facial features within the portion of the face. The systems and methods generate a descriptor function representing the set of facial features, fit an objective function of the descriptor function, identify a recognition probability for each facial feature, and assign an identification to each facial feature.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 201780064382.6, which was filed on October 19, 2017, had a priority date of October 19, 2016, entered the Chinese national phase on April 18, 2019, and is entitled "Method and System for Facial Modeling".

[0002] Priority requirements

[0003] This application claims priority to U.S. Application No. 15 / 297,789, filed October 19, 2016, the entire contents of which are incorporated herein by reference. Technical Field

[0004] Embodiments of this disclosure generally relate to the automatic processing of images. More specifically, but not as a limitation, this disclosure proposes systems and methods for modeling representations of faces depicted within a set of images. Background Technology

[0005] Telecommunications applications and devices can use various media, such as text, images, audio recordings, and / or video recordings, to provide communication between multiple users. For example, video conferencing allows two or more individuals to communicate with each other using a combination of software applications, telecommunications devices, and telecommunications networks. Telecommunications devices can also record video streams for transmission as messages over telecommunications networks.

[0006] Currently, facial processing technologies used for communication or recognition purposes are typically user-selectable and guided. Facial recognition technologies usually train models for individual features, such that training a first model for the first feature appearing on the face is performed separately from training a second model for the second feature appearing on the face. When modeling a new face or performing a recognition function, the separately trained models are typically used individually and sequentially to build the model of the new face or for recognition. Attached Figure Description

[0007] The accompanying drawings illustrate only exemplary embodiments of this disclosure and should not be construed as limiting its scope.

[0008] Figure 1 This is a block diagram illustrating a networking system according to some example embodiments.

[0009] Figure 2 This is a diagram illustrating a facial modeling system according to some example embodiments.

[0010] Figure 3 This is a flowchart illustrating an example method for modeling and recognizing various aspects of a face from a set of images, according to some example embodiments.

[0011] Figure 4This is a flowchart illustrating an example method for modeling and recognizing various aspects of a face from a set of images, according to some example embodiments.

[0012] Figure 5 This is a flowchart illustrating an example method for modeling and recognizing various aspects of a face from a set of images, according to some example embodiments.

[0013] Figure 6 This is a flowchart illustrating an example method for modeling and recognizing various aspects of a face from a set of images, according to some example embodiments.

[0014] Figure 7 This is a flowchart illustrating an example method for modeling and recognizing various aspects of a face from a set of images, according to some example embodiments.

[0015] Figure 8 This is a flowchart illustrating an example method for modeling and recognizing various aspects of a face from a set of images, according to some example embodiments.

[0016] Figure 9 These are user interface diagrams depicting example mobile devices and mobile operating system interfaces according to some example embodiments.

[0017] Figure 10 This is a block diagram illustrating an example of a software architecture that can be installed on a machine according to some example embodiments.

[0018] Figure 11 It is a block diagram that presents a graphical representation of a machine in the form of a computer system according to an example embodiment, within which a set of instructions can be executed to cause the machine to perform any of the methods discussed herein.

[0019] The titles provided in this article are for convenience only and do not necessarily affect the scope or meaning of the terms used. Detailed Implementation

[0020] The following description includes systems, methods, techniques, instruction sequences, and computer program products illustrating embodiments of the present disclosure. In this description, numerous specific details are set forth for purposes of explanation in order to provide an understanding of various embodiments of the subject matter of the invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the invention can also be practiced without these specific details. Generally, well-known examples of instructions, protocols, structures, and techniques need not be shown in detail.

[0021] While methods exist for modeling faces and facial features within images, these existing methods do not use neural models to model visual representations. Furthermore, these methods do not employ unified models to simultaneously model, fit, or recognize multiple facial features or attributes. Existing methods typically train each feature or aspect independently, or use different and independent tasks as regularizers to train each feature or aspect. Unified models can be based on a unified function with an object function that is simultaneously trained and modified (e.g., theoretically optimized) to interpret aspects of the face using neural models based on low-level curvature and other visual representations. Unified models employ a unified deep neural network model to simultaneously train all facial attributes. Unified models can achieve shared representations for multiple attribute recognition tasks. Therefore, there remains a need in the art to improve the identification, modeling, interpretation, and recognition of faces within images with minimal or no user interaction. Furthermore, there remains a need in the art to improve the generation of facial models and the recognition of face-related interpretive or inferential aspects (characteristics not directly identifiable on the face). As described herein, methods and systems are presented for modeling and recognizing facial features or interpreting aspects of a face based on facial features or facial landmarks depicted within an image using a single user interaction of initial selection.

[0022] Embodiments of this disclosure may generally relate to automated image segmentation and generation of facial representations or models based on neural network processing of segmented images and features identified within the images. In one embodiment, a facial modeling system accesses or receives an image depicting a face. The facial modeling system identifies facial features and facial landmarks depicted within the image. Facial features and facial landmarks are modeled within a single function by fitting each facial feature in combination with other facial features to draw conclusions based on previously trained facial data. The facial modeling system generates probabilities indicative of information about the face within the image based on the function and the interrelationships between its components. In some cases, the facial modeling system combines normalized features from an average face, which are combined to form a single aggregated face and identify information including demographic information.

[0023] The above is a specific example. Various embodiments of this disclosure relate to instructions for apparatuses and one or more processors of the apparatus to perform automatic inference (e.g., real-time modification of the video stream) on faces within an image or video stream transmitted from one apparatus to another apparatus during video stream acquisition. A face modeling system is described that identifies and generates inferences about objects and regions of interest within an image or between video streams, and through a set of images including the video stream. In various example embodiments, the face modeling system identifies and tracks one or more facial features depicted in the video stream or within an image, and performs image recognition, face recognition, and face processing functions regarding one or more facial features, as well as correlations between two or more facial features.

[0024] Figure 1 This is a network diagram depicting a network system 100 according to one embodiment, having a client-server architecture configured for exchanging data over a network. For example, network system 100 may be a messaging system in which clients transmit and exchange data within network system 100. The data may relate to various functions (e.g., sending and receiving text and media communications, determining geographic location, etc.) and aspects associated with network system 100 and its users (e.g., transmitting communication data, receiving and sending indications of communication sessions, etc.). While network system 100 is shown herein as a client-server architecture, other embodiments may include other network architectures, such as peer-to-peer or distributed network environments.

[0025] like Figure 1 As shown, network system 100 includes social messaging system 130. Social messaging system 130 is typically based on a three-tier architecture, comprising an interface layer 124, an application logic layer 126, and a data layer 128. As understood by those skilled in the art of computers and the Internet, Figure 1 Each component or engine shown represents a set of executable software instructions and corresponding hardware (e.g., memory and processor) for executing the instructions, forming a hardware-implemented component or engine, and serving as a dedicated machine configured to perform a specific set of functions when executing the instructions. To avoid obscuring the subject matter of the invention with unnecessary detail, from Figure 1 Various functional components and engines that are not closely related to conveying an understanding of the subject matter of this invention have been omitted. Of course, additional functional components and engines can be integrated with social messaging systems (such as...) Figure 1 This can be used in conjunction with the social messaging system shown to enable additional functionalities not specifically described herein. Furthermore, Figure 1 The various functional components and engines described can reside on a single server computer or client device, or they can be distributed across several server computers or client devices in various arrangements. Furthermore, although... Figure 1 The social messaging system 130 is described as having a three-layer architecture, but the subject matter of this invention is by no means limited to this architecture.

[0026] like Figure 1As shown, interface layer 124 includes interface component (e.g., web server) 140, which receives requests from various client computing devices and servers, such as client device 110 executing client application 112 and third-party server 120 executing third-party application 122. In response to a received request, interface component 140 transmits an appropriate response to the requesting device via network 104. For example, interface component 140 may receive requests such as Hypertext Transfer Protocol (HTTP) requests or other web-based application programming interface (API) requests.

[0027] Client device 110 can execute a conventional web browser application or an application (also referred to as an "app") that has been developed for a specific platform to include any of a variety of mobile computing devices and mobile-specific operating systems (e.g., iOS™, Android™, Windows® Phone). Furthermore, in some example embodiments, client device 110 forms all or part of facial modeling system 160, such that components of facial modeling system 160 configure client device 110 to perform a specific set of functions relating to the operation of facial modeling system 160.

[0028] In the example, client device 110 executes client application 112. Client application 112 can provide the functionality to present information to user 106 and to communicate via network 104 to exchange information with social messaging system 130. Furthermore, in some examples, client device 110 performs the functions of face modeling system 160 to segment images of the video stream during video stream acquisition and transmit the video stream (e.g., using image data modified based on the segmented images from the video stream).

[0029] Each of the client devices 110 may include a computing device, which includes at least a display and the ability to communicate with the network 104 to access the social messaging system 130, other client devices, and the third-party server 120. Client devices 110 include, but are not limited to, remote devices, workstations, computers, general-purpose computers, internet devices, handheld devices, wireless devices, portable devices, wearable computers, cellular or mobile phones, personal digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, desktop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, network PCs, minicomputers, etc. User 106 may be a person, a machine, or other component that interacts with client device 110. In some embodiments, user 106 interacts with the social messaging system 130 via client device 110. User 106 may not be part of a networked environment but may be associated with client device 110.

[0030] like Figure 1As shown, data layer 128 has a database server 132 that facilitates access to an information repository or database 134. Database 134 is a storage device for storing data such as member profile data, social graph data (e.g., relationships between members of the social messaging system 130), image modification preference data, accessibility data, and other user data.

[0031] Individuals can register with the social messaging system 130 to become members of the social messaging system 130. Once registered, members can form social network relationships (e.g., friends, followers, or contacts) on the social messaging system 130 and interact with a wide range of applications offered by the social messaging system 130.

[0032] Application logic layer 126 includes various application logic components 150, which, in conjunction with interface component 140, use data obtained from various data sources or data services in data layer 128 to generate various user interfaces. Each application logic component 150 can be used to implement functionality associated with various applications, services, and features of social messaging system 130. For example, a social messaging application can be implemented together with one or more application logic components 150. The social messaging application provides a messaging mechanism for users of client device 110 to send and receive messages including text and media content such as pictures and videos. Client device 110 can access and view messages from the social messaging application for a specified time period (e.g., limited or unlimited). In the example, a message recipient can access a specific message for a predefined duration (e.g., specified by the message sender), which begins when the specific message is first accessed. After the predefined duration has elapsed, the message is deleted, and the message recipient can no longer access the message. Of course, other applications and services can be embodied in their respective application logic components 150.

[0033] like Figure 1 As shown, the social messaging system 130 may include at least a portion of the facial modeling system 160, which is capable of recognizing, tracking, and modifying video data during video data acquisition by the client device 110. Similarly, as described above, the client device 110 includes a portion of the facial modeling system 160. In other examples, the client device 110 may include the entire facial modeling system 160. Where the client device 110 includes a portion (or all) of the facial modeling system 160, the client device 110 may operate independently or collaboratively with the social messaging system 130 to provide the functionality of the facial modeling system 160 described herein.

[0034] In some embodiments, the social messaging system 130 may be a short-time messaging system that allows short-term communication in which content (e.g., video clips or images) is deleted after a deletion trigger event, such as viewing time or viewing completion. In this embodiment, the apparatus uses various components described herein in any aspect of generating, sending, receiving, or displaying short-time messages. For example, the apparatus implementing the facial modeling system 160 can identify, track, and modify objects of interest, such as pixels representing skin on a face depicted in a video clip. The apparatus can modify the objects of interest during video clip acquisition as part of the content of a short-time message without requiring image processing after the video clip is acquired.

[0035] exist Figure 2 In various embodiments, the facial modeling system 160 may be implemented as a standalone system or in conjunction with the client device 110, and is not necessarily included in the social messaging system 130. The facial modeling system 160 is shown as including an access component 210, a recognition component 220, a facial processing component 230, a description component 240, a modeling component 250, a probability component 260, an allocation component 270, and a mapping component 280. All or some of the components 210-280 communicate with each other, for example via network coupling, shared memory, etc. Each component of 210-280 may be implemented as a single component, combined with other components, or further subdivided into multiple components. Other components unrelated to the example embodiment may also be included, but are not shown.

[0036] Access component 210 accesses or otherwise acquires images captured by an image acquisition device or otherwise received by or stored in client device 110. In some cases, access component 210 may include part or all of an image acquisition component configured to enable the image acquisition device of client device 110 to acquire images based on user interaction with a user interface presented on the display device of client device 110. Access component 210 may transfer images or portions of images to one or more other components of the face modeling system 160.

[0037] The recognition component 220 identifies faces or other regions of interest within an image or set of images received from the access component 210. In some embodiments, the recognition component 220 tracks the identified faces or regions of interest across multiple images in a set of images (e.g., a video stream). The recognition component 220 may pass values ​​representing faces or regions of interest (e.g., coordinates within an image or a portion of an image) to one or more components of the face modeling system 160.

[0038] The face processing component 230 identifies facial features or facial landmarks depicted on the face or within a region of interest identified by the recognition component 220. In some embodiments, the face processing component 230 identifies, in addition to facial landmarks depicted on the face or within a region of interest, facial landmarks that are expected to be present but are missing. The face processing component 230 can determine the orientation of the face based on the facial landmarks and can identify one or more relationships between facial landmarks. The face processing component 230 can pass values ​​representing facial landmarks to one or more components of the face modeling system 160.

[0039] The description component 240 generates descriptor functions. In some cases, a set of object functions is used to generate the descriptor functions. The object functions may correspond to individual facial features. In some embodiments, the object functions correspond to feature extraction and classification of the extracted features. The descriptor functions may be generated during the training process and may be accessed by the description component 240 during experimental use or during in-app deployment on a client device.

[0040] Modeling component 250 fits each object function to the descriptor function. Modeling component 250 may use stochastic gradient descent to fit the object functions, as described below. In some embodiments, modeling component 250 simultaneously modifies the parameters of each object function. Modeling component 250 may modify the parameters of the object functions relative to a regularization function or by modifying the parameters of the regularization function. Modeling component 250 may determine the regularization function during the fitting process based on the expected value compared to the floating output of the descriptor function or the object function.

[0041] The probability component 260 identifies the recognition probability of each facial feature. The recognition probability can be provided by the descriptor function in response to image processing. In some cases, the recognition probability is a value representing the probability that a facial feature corresponds to a specified characteristic. The probability component 260 can identify the numerical value associated with the recognition probability and provide the value to the allocation component 270.

[0042] The assignment component 270 assigns an identifier to each facial feature regarding the probability of recognition for that facial feature. The assignment component 270 may assign identifiers to all features based on an identifier level (e.g., a broad identifier such as male / female). In some embodiments, the assignment component 270 generates a notification instructing the assignment of an identifier to one or more facial features.

[0043] Mapping component 280 maps a face received from a client device to a reference face. Mapping component 280 can map the face and reference face using facial features or landmarks of the face and reference facial landmarks of the reference face. In some embodiments, mapping component 280 can modify the image of the face to perform the mapping without introducing skew or other ratio changes to the face. In some cases, mapping component 280 performs the mapping by prioritizing the alignment of one or more facial landmarks or by aligning a maximum number of facial landmarks.

[0044] Figure 3 A flowchart illustrating an example method 300 for modeling and recognizing various aspects of a face from a set of images (e.g., a video stream) is provided. The operation of method 300 can be performed by components of a face modeling system 160, and is described below for illustrative purposes.

[0045] In operation 310, access component 210 receives or otherwise accesses one or more images depicting at least a portion of a face. In some embodiments, access component 210 receives one or more images, such as frames of a video stream captured by an image acquisition device associated with client device 110. In some cases, the video stream is presented on a user interface of a face modeling application. Access component 210 may include an image acquisition device as part of the hardware including access component 210. In these embodiments, access component 210 directly receives one or more images or video streams captured by the image acquisition device. In some cases, access component 210 passes all or part of one or more images or video streams (e.g., a set of images including a video stream) to one or more components of face modeling system 160, as described in more detail below.

[0046] In operation 320, the recognition component 220 detects portions of a face depicted within one or more images. In some embodiments, the recognition component 220 includes a set of face tracking operations to identify faces or portions of faces within one or more images. The recognition component 220 may use the Viola-Jones object detection framework, intrinsic surface techniques, genetic algorithms for face detection, edge detection methods, or any other suitable object class detection methods or set of operations to identify faces or portions of faces within one or more images. Where one or more images are multiple images (e.g., a set of images in a video stream), after identifying faces or portions of faces in the initial image, the face tracking operations of the recognition component 220 may identify changes in the position of faces across multiple images, thereby tracking facial movement within multiple images. Although specific techniques have been described, it should be understood that the recognition component 220 may use any suitable technique or set of operations to identify faces or portions of faces within one or more images without departing from the scope of this disclosure.

[0047] In operation 330, face processing component 230 identifies a set of facial features within a portion of a face depicted in one or more images. The facial features may correspond to one or more facial landmarks identified by face modeling system 160. Facial features can be identified in response to detecting a face in one or more faces. In some embodiments, face processing component 230 identifies a set of facial features within a portion of a face in a subset of one or more images. For example, face processing component 230 may identify a set of facial features in a set of images (e.g., a first set of images) of multiple images, wherein portions of the face or facial landmarks appear in that set of images but not in the remaining images of the multiple images (e.g., a second set of images). In some embodiments, the identification of facial features or facial landmarks may be performed using a face tracking operation in conjunction with the detection operations described above, as a sub-operation or part of the identification of a face or a portion of a face.

[0048] In operation 340, the description component 240 generates a descriptor function. The descriptor function may be generated based on a set of facial features identified in relation to operation 330. In some example embodiments, the descriptor function represents the set of facial features and includes a set of object functions. Each object function represents a facial feature, a recognition characteristic, a set of recognition characteristics, a classification characteristic, a category, or any other suitable aspect capable of recognizing, characterizing, or classifying a face. The descriptor function and the set of object functions can represent the neural network model as a deep neural network structure. The neural network structure may include a varying number of layers (e.g., object functions). The number and type of layers (e.g., object functions) may vary based on the amount and type of information to be interpreted or otherwise recognized for the face. In some embodiments, the layers include one or more convolutional layers, one or more pooling layers, and one or more fully connected layers.

[0049] In operation 350, modeling component 250 fits each of a set of object functions. In some embodiments, the object functions are fitted in response to generating a descriptor function. The object functions can be fitted simultaneously and relative to each other, such that a modification to one object function in the set causes a corresponding modification to another object function. For example, in the case where a first object function identifies the race of a face in one or more images, a second object function identifies the gender of the face, and a third object function identifies the age of the face, the first, second, and third object functions can establish or inform each other. The first object function identifying race can correspond to specified predetermined parameters of features identified on the face. For example, the first object function identifying one or more races can identify a set of low-level curvatures or other visual representations associated with wrinkle patterns associated with certain ages or with bimorphic characteristics associated with the gender of the face. In some cases, modeling component 250 fits object functions in a cascaded manner, wherein each fitted object function causes a modification to the parameters of one or more subsequent object functions within the descriptor function.

[0050] In some embodiments, the object function corresponds to low-level patterns or curvatures in the face, rather than specifically to a particular feature or characteristic of the face being recognized. In these cases, the modeling component 250 can simultaneously fit each object function such that determining the fit for each object function generates a set of probabilities for a predetermined set of features or recognition characteristics, as the output of the descriptor function.

[0051] In operation 360, the probability component 260 identifies the recognition probability of each facial feature. The recognition probability can be a value representing the probability that a facial feature corresponds to a specified characteristic. In some embodiments, the object function corresponds to one or more recognition probabilities of the facial feature. In some embodiments, the recognition probability is a bounded probability, such that the recognition probability is contained within desired maximum and minimum boundaries. Within these maximum and minimum boundaries, there may be two or more feature values ​​corresponding to the identified feature, characteristic, demographic, or other facial recognition aspect. For example, feature values ​​may correspond to an age range, one or more ethnicities, one or more biological sexes, or any other suitable facial recognition aspect. Where the recognition probability is defined between two feature values ​​(such as values ​​corresponding to male and female), the probability component 260 can identify values ​​falling at or between a first value associated with male and a second value associated with female. The proximity of the recognition probability value to one of the first or second values ​​indicates the probability that the facial feature corresponds to male or female, respectively.

[0052] In operation 370, the assignment component 270 assigns an identifier to each facial feature based on the recognition probability for each facial feature recognition. In some embodiments, the assignment component 270 assigns identifiers to all facial features for which a recognition probability has been identified. For example, in the case where it is determined that the face is of male gender, the assignment component 270 may assign each facial feature identified on the face to male. When the assignment component 270 assigns identifiers such as gender to facial features in a set of training data and the assigned identifiers are correct, the low-level curvature or other visual representations associated with these facial features are identified as corresponding to the assigned gender and can be used to inform the probability calculation of gender in subsequent datasets.

[0053] In cases where the assignment component 270 assigns an identifier to facial features in a set of experimental or user-provided data, the assignment component 270 may generate a notification of the identifier and cause the notification to be displayed on a display device of the client device. For example, in cases where a user of the client device causes the acquisition or access to an image of the user's face, the facial modeling system 160 (including the assignment component 270) may determine identification information about the user, such as age range, ethnicity, gender, and other identification data. The assignment component 270 may generate graphical interface elements, such as tables, lists, or other graphical representations of one or more of the user's age, ethnicity, gender, and other identification information. In some cases, the assignment component 270 includes an indication of the facial features associated with each identification information element provided in the notification. The assignment component 270 may also include a probability value for each facial feature to indicate the impact of each facial feature on the determination and assignment of identification information.

[0054] In some embodiments, the notification is provided in the form of an automatically generated avatar. The modeling component 250, in conjunction with the allocation component 270, can identify facial features of a user-provided face within a library of modeled facial features and generate a graphical representation of the user-provided face. The graphical representation of the face can be a cartoon or other artistic representation, animation, or any other suitable avatar. Modeled facial features can be selected and rendered on the avatar based on a determined similarity between the modeled facial features and the facial features of the user-provided face. In some cases, the modeled facial features include metadata. The metadata may include one or more of recognition probabilities, feature values, and recognition information (e.g., demographic information matching the modeled facial features). In some cases, the modeling component 250 and the allocation component 270 may use the metadata to select modeled facial features by matching the metadata with a recognition probability or assigned identifier determined by the probability component 260 and the allocation component 270 accordingly.

[0055] Figure 4A flowchart illustrating an example method 400 for modeling and recognizing various aspects of a face from a set of images is shown. The operations of method 400 can be performed by components of the face modeling system 160. In some cases, certain operations of method 400 can be performed using one or more operations of method 300, or as sub-operations of one or more operations of method 300, as will be explained in more detail below.

[0056] In operation 410, access component 210 accesses a reference image from a facial reference database. The reference image has a set of reference facial landmarks. In some embodiments, operation 410 is performed in response to operation 310 or as a sub-operation of operation 310. The reference image may be represented as a normalized representation of facial size, position, and orientation within a frame of a video stream or image. In some cases, the reference image includes a reference face or a normalized face. The reference face may include the average of faces used to generate the training model of the facial modeling system 160. In some embodiments, the reference image is a synthetic face representing a normalized face generated from multiple faces.

[0057] In operation 420, mapping component 280 maps a face and a set of facial features to a reference image. Mapping component 280 can map a face received in the image of operation 310 to a reference face in the reference image. In some embodiments, mapping component 280 maps a face to a reference face by aligning or substantially aligning facial features or facial landmarks identified in operation 330 with reference facial landmarks of the reference face. Mapping component 280 can modify the size, orientation, and position of the face to map facial landmarks of the face to reference facial landmarks. In some cases, mapping component 280 modifies the size, orientation, and position without modifying the aspect ratio or introducing skew or other deformities into the face.

[0058] Once the mapping component 280 modifies the size, orientation, and position of the face to approximate a reference face, the mapping component 280 can align one or more facial landmarks of the face with reference facial landmarks. If a facial landmark is not precisely aligned with a reference facial landmark, the mapping component 280 can position the facial landmark close to the reference facial landmark. The mapping component 280 can align the face and the reference face by positioning the face such that a maximum number of facial landmarks overlap with the reference facial landmark. The mapping component 280 can also align the face and the reference face by positioning facial landmarks to overlap with reference facial landmarks that have been identified (e.g., previously determined) and have an importance value higher than a predetermined importance threshold.

[0059] Figure 5A flowchart illustrating an example method 500 for modeling and recognizing various aspects of a face from a set of images is depicted. Operations of method 500 can be performed by components of the face modeling system 160. In some cases, in one or more of the described embodiments, certain operations of method 500 can be performed using one or more operations of method 300 or 400, or as sub-operations of one or more operations of method 300 or 400, as will be explained in more detail below.

[0060] In operation 510, modeling component 250 modifies one or more object functions within the descriptor function to fit each object function of the descriptor function via stochastic gradient descent updating. In some embodiments, the stochastic gradient descent algorithm optimizes (e.g., attempts to identify the theoretical optimum) the object functions. Modeling component 250 employing stochastic gradients can perform gradient updates to fit the function collaboratively. Operation 510 can be performed in response to operation 350 described above, or as a sub-operation or part of operation 350.

[0061] In some embodiments, stochastic gradient descent can be performed by generating a floating output for one or more object functions or descriptor functions. Modeling component 250 identifies at least one recognition probability generated within the floating output. For each recognition probability, modeling component 250 determines a residual between the desired feature value and the recognition probability value. For example, in the case of a recognition probability of 0.7 and a desired feature value of 1.0 (indicating male), the residual is 0.3. Modeling component 250 may propagate the residual (e.g., a loss between the recognition value and the desired value). Modeling component 250 may modify one or more object functions or parameters associated with object functions to generate a recognition probability of 1.0 or closer to 1.0 than the initially provided value of 0.7. In the absence of a single underlying fact or a value for both desired and affirmative recognition, modeling component 250 may determine the desired value based on one or more training or reference images.

[0062] In operation 520, modeling component 250 selects a first object function as the regularization function. In some cases, the first object function is naturally used as the regularization function. The natural regularization function can be associated with a recognition probability having a verifiable underlying fact or expected value. For example, the regularization function could be gender, where gender is known, and therefore the expected value of stochastic gradient descent is known. Modeling component 250 can select the first object based on the known recognition probability and the existence of the expected value. When two or more object functions are associated with the known and expected values ​​of the recognition probability, modeling component 250 can select the object function associated with the recognition probability that provides a floating output closer to the expected value as the regularization function.

[0063] In operation 530, modeling component 250 modifies one or more of the remaining object functions of a plurality of object functions relative to a regularization function. In some embodiments, modeling component 250 modifies the plurality of object functions by modifying the parameters associated with each function. Modeling component 250 may iteratively modify the parameters of all object functions (including the regularization function) until the regularization function provides a floating output value of a recognition probability that matches the expected value or is within a predetermined error threshold of the expected value.

[0064] Figure 6 A flowchart illustrating an example method 600 for modeling and recognizing various aspects of a face from a set of images is shown. Operations of method 600 can be performed by components of the face modeling system 160. In some cases, certain operations of method 600 can be performed using one or more operations of methods 300, 400, or 500, or as sub-operations of one or more operations of methods 300, 400, or 500, as will be explained in more detail below.

[0065] In operation 610, modeling component 250 divides one or more images into a set of pixel regions. In these embodiments, the set of object functions may include functions of different types as part of a descriptor function. In some cases, the first specified function is a convolutional layer describing at least one facial feature of a set of facial features. One or more images may be represented as a two-dimensional matrix consisting of a set of pixel regions.

[0066] In some example embodiments, convolutional layers can be used as feature descriptors, either alone or in combination with another layer. The feature descriptors identify unique aspects or combinations of aspects on a face to distinguish the face based on at least a portion of a reference face. In some cases, convolutional layers (e.g., a set of object functions) can be represented using a set of matrices. A convolutional layer can be represented as described below with respect to Equation 1.

[0067]

[0068] Equation 1

[0069] As shown in Equation 1, two matrices A and B can be represented by M. A xN A and M B xN B Presented. Matrix C can be calculated according to Equation 1 above. Furthermore, in Equation 1, as well as .

[0070] In operation 620, modeling component 250 performs a first specified function (e.g., a convolutional layer) over a set of pixel regions. In some embodiments, modeling component 250 performs the specified function over each pixel region of the set of pixel regions. The convolutional layer can act as a filter operation. Modeling component 250 can process the image or each pixel in each pixel region using the filter. In some embodiments, the size of the filter can be determined based on the size or measurement of the image. The filter can provide a single response for each location. Processing each pixel or each pixel region can be performed by: identifying one or more values ​​of the pixel or pixel region, extracting a region around the pixel or pixel region of a size determined according to the filter, and generating a single output for the pixel or pixel region based on the values ​​of the extracted region. The extracted region can be determined by one or more of the pixel region size or the image size. For example, in the case where the pixel region is a single pixel, the filter size can be a multiple of the pixel region, such as 2 times. In this example, for each pixel, the extracted region can be a two-pixel multiplied by two-pixel matrix surrounding the filtered pixel.

[0071] In operation 630, modeling component 250 generates values ​​representing the presence of visual features within pixel regions. In some embodiments, modeling component 250 generates values ​​for each pixel region in a set of pixel regions. Modeling component 250 may generate values ​​for pixel regions in response to executing a first specified function on each pixel region. The values ​​of pixel regions may represent the presence of visual features, wherein a defined value for a pixel region exceeds a predetermined threshold.

[0072] In operation 640, modeling component 250 generates a result matrix that includes values ​​representing the presence of visual features within pixel regions. The result matrix can be a representation of the image after processing by the first function. For example, the image can be divided into pixel regions, and the value of each pixel region obtained by the first object function can be the value of that pixel region in the result matrix.

[0073] Figure 7 A flowchart illustrating an example method 700 for modeling and recognizing various aspects of a face from a set of images is shown. Operations of method 700 can be performed by components of the face modeling system 160. In some cases, certain operations of method 700 can be performed using one or more operations of methods 300, 400, 500, or 600, or as sub-operations of one or more operations of methods 300, 400, 500, or 600, as will be explained in more detail below.

[0074] In operation 710, modeling component 250 identifies the maximum value of one or more sub-regions of each pixel region. In some embodiments, operation 710 may be performed in response to one or more images being divided into a set of pixel regions as in operation 610, and as in operation 640, generating a result matrix. If modeling component 250 has generated a result matrix, it may identify the maximum value of the sub-regions of each pixel region by executing a second specified function. The second specified function may be a pooling layer configured to reduce the size of the result matrix. In some embodiments, pooling layers are used as feature descriptors, alone or in combination with other layers. Pooling layers may reduce the size of the output of another layer (e.g., a convolutional layer). In some embodiments, pooling layers include max pooling layers and average pooling layers. For example, an initial pooling layer may contain 4x4 elements in its element matrix. When a max pooling function is applied to a pooling layer, the pooling layer generates a subset matrix. For example, a pooling layer may reduce a 4x4 element matrix by generating a 2x2 matrix. In some embodiments, the pooling layer outputs a representative value for each group of elements within a 4x4 matrix (e.g., four groups of elements).

[0075] In operation 720, modeling component 250 selects the maximum value of each sub-region as the representation of the pixel region that includes that sub-region in the resulting matrix. For example, a pixel region can be divided into four independent elements of the resulting matrix, each element having a different value. Modeling component 250 selects the value of an element within the pixel region that is greater than the remaining values ​​of the four independent elements.

[0076] In operation 730, modeling component 250 generates a pooling matrix that includes the maximum value of each pixel region. The pooling matrix can represent a reduction in the size of the resulting matrix. For example, if the resulting matrix includes a set of two-by-two elements for each pixel region and comprises four total pixel regions, the pooling matrix can be a two-by-two matrix, thus reducing the overall size of the resulting matrix.

[0077] Figure 8 A flowchart illustrating an example method 800 for modeling and recognizing various aspects of a face from a set of images is shown. Operations of method 800 can be performed by components of the face modeling system 160. In some cases, certain operations of method 800 can be performed using one or more operations of methods 300, 400, 500, 600, or 700, or as sub-operations of one or more operations of methods 300, 400, 500, 600, or 700, as will be explained in more detail below.

[0078] In operation 810, modeling component 250 generates an inner product of the pooling matrix and layer parameters. This inner product can be generated by a fully connected layer. A fully connected layer can represent each node in the layer as connected to all nodes from previous layers. For example, a specified region or element within a fully connected layer can include connections between the pooling matrix and layer parameters. In some embodiments, the output of a fully connected layer can be the inner product of the outputs of previous layers.

[0079] In some embodiments, fully connected layers can operate as classifiers to generalize common patterns from a large set of images. For example, given labels and images of men and women, fully connected layers can be trained to modify the parameter values ​​of each object function of the descriptor function to correctly identify the labels when processing the images using the descriptor function. In these cases, convolutional and pooling layers can identify and extract facial features and provide feature vectors. Fully connected layers can classify images based on one or more generated recognition probabilities.

[0080] In operation 820, modeling component 250 generates recognition probabilities based on inner product, pooling matrix, and result matrix. In some embodiments, the recognition probability is a probability value defined between a first value and a second value. The first value and the second value may represent a binary option represented by the probability value. The recognition probability value may fall between the first value and the second value, indicating the likelihood that a facial feature or aspect corresponds to a first feature represented by the first value or a second feature represented by the second value. In some cases, the first value and the second value may represent boundary values, where one or more additional values ​​are located between the first value and the second value. In these cases, the recognition probability value falling between the first value and the second value may represent the likelihood that a facial feature or aspect corresponds to a characteristic represented by a value including a combination of the first value, one or more additional values, and the second value.

[0081] In some cases, the facial modeling system 160 identifies gender features corresponding to a first value and a second value. In some embodiments, the modeling component 250 performs operation 822 to identify a gender threshold between the first value and the second value. The gender threshold can be located equidistant from the first value and the second value. In some cases, based on modifications to the descriptor function and object function described above with respect to method 500, the gender threshold can be located to a value closer to either the first value or the second value.

[0082] In operation 824, modeling component 250, alone or in combination with allocation component 270, selects the gender for a face based on the recognition probability value exceeding a gender threshold, either a first value or a second value. Once modeling component 250 or allocation component 270 selects the gender, allocation component 270 can generate a notification as described in operation 370 above.

[0083] Modules, components and logic

[0084] Certain embodiments herein are described as including logic or multiple components, modules, or mechanisms. Components may constitute hardware components. A “hardware component” is a tangible unit capable of performing certain operations and may be configured or arranged in some physical manner. In various example embodiments, a computer system (e.g., a standalone computer system, a client computer system, or a server computer system) or a hardware component of a computer system (e.g., at least one hardware processor, processor, or a group of processors) is configured by software (e.g., an application or an application portion) to perform certain operations as described herein.

[0085] In some embodiments, the hardware component is implemented mechanically, electronically, or in any suitable combination thereof. For example, the hardware component may include dedicated circuitry or logic permanently configured to perform certain operations. For instance, the hardware component may be a dedicated processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). The hardware component may also include programmable logic or circuitry temporarily configured by software to perform certain operations. For instance, the hardware component may include software contained within a general-purpose processor or other programmable processor. It is understood that cost and time considerations may drive the decision to implement the hardware component mechanically in dedicated and permanently configured circuitry or in temporarily configured circuitry (e.g., configured by software).

[0086] Therefore, the phrase "hardware component" should be understood to include tangible entities, which are physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate or perform certain operations described herein in a certain way. As used herein, "hardware-implemented component" refers to a hardware component. Considering embodiments in which hardware components are temporarily configured (e.g., programmed), it is not necessary to configure or instantiate every hardware component at any given time. For example, where the hardware components include a general-purpose processor configured by software as a dedicated processor, the general-purpose processor can be configured at different times as correspondingly different dedicated processors (e.g., including different hardware components). Thus, software can configure a particular one or more processors, for example, constituting a particular hardware component at one time and different hardware components at different times.

[0087] Hardware components can provide and receive information from other hardware components. Therefore, the described hardware components can be considered communicatively coupled. In the presence of multiple hardware components, communication can be achieved through signal transmission (e.g., via appropriate circuitry and buses) between or among two or more hardware components. In embodiments where multiple hardware components are configured or instantiated at different times, communication between these hardware components can be achieved, for example, by storing and retrieving information from memory structures accessible to the multiple hardware components. For example, one hardware component performs an operation and stores the output of that operation in a memory device communicatively coupled to it. Another hardware component can then later access the memory device to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and can operate on resources (e.g., information sets).

[0088] The various operations of the example methods described herein can be performed, at least in part, by a processor configured either temporarily (e.g., by software) or permanently to perform the relevant operations. Whether temporarily or permanently configured, such a processor constitutes a component of a processor implementation that operates to perform the operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using a processor.

[0089] Similarly, the methods described herein can be implemented at least in part by a processor, where a particular processor or processors are examples of hardware. For example, at least some operations of the method can be performed by a processor or a component implemented by a processor. Furthermore, the processor can also operate to support the performance of the relevant operations in a “cloud computing” environment or as “Software as a Service” (SaaS). For example, at least some operations can be performed by a group of computers (as an example of a machine including a processor), which can be accessed via a network (e.g., the Internet) and via a suitable interface (e.g., an application programming interface (API)).

[0090] The performance of certain operations can be distributed among processors, residing not only within a single machine but also deployed across multiple machines. In some example embodiments, the processor or processor-implemented component resides in a single geographic location (e.g., within a home environment, office environment, or server cluster). In other example embodiments, the processor or processor-implemented component is distributed across multiple geographic locations.

[0091] application

[0092] Figure 9An example mobile device 900 executing a mobile operating system (e.g., iOS™, Android™, Windows® Phone, or other mobile operating systems) is shown, consistent with some embodiments. In one embodiment, the mobile device 900 includes a touchscreen operable to receive tactile data from a user 902. For example, the user 902 may physically touch 904 of the mobile device 900, and in response to the touch 904, the mobile device 900 may determine tactile data, such as touch location, touch force, or gesture action. In various example embodiments, the mobile device 900 displays a home screen 906 (e.g., Springboard on iOS™), operable to launch applications or otherwise manage various aspects of the mobile device 900. In some example embodiments, the home screen 906 provides status information such as battery life, connectivity, or other hardware status. The user 902 can activate user interface elements by touching areas occupied by corresponding user interface elements. In this way, the user 902 interacts with applications on the mobile device 900. For example, touching an area occupied by a specific icon included in the home screen 906 results in launching an application corresponding to that specific icon.

[0093] like Figure 9 As shown, the mobile device 900 may include an imaging device 908. The imaging device 908 may be a camera or any other device coupled to the mobile device 900 capable of acquiring a video stream or one or more consecutive images. The imaging device 908 may be triggered by the face modeling system 160 or optional user interface elements to initiate the acquisition of a continuum of video streams or images and to pass the continuum of video streams or images to the face modeling system 160 for processing according to one or more methods described in this disclosure.

[0094] Many kinds of applications (also referred to as "application software") can be executed on mobile device 900, such as native applications (e.g., applications programmed in Objective-C, Swift, or another suitable language running on iOS™, or applications programmed in Java running on Android™), mobile web applications (e.g., applications written in Hypertext Markup Language-5 (HTML5)), or hybrid applications (e.g., native shell applications that launch HTML5 sessions). For example, mobile device 900 includes messaging applications, audio recording applications, camera applications, book reader applications, media applications, fitness applications, file management applications, location applications, browser applications, settings applications, contact applications, telephone calling applications, or other applications (e.g., game applications, social networking applications, biometric monitoring applications). In another example, mobile device 900 includes social messaging application software 910 such as SNAPCHAT®, which, consistent with some embodiments, allows users to exchange short messages including media content. In this example, social messaging application software 910 may incorporate aspects of the embodiments described herein. For example, in some embodiments, the social messaging application 910 includes short-lived media galleries created by the user's social messaging application 910. These galleries may include videos or pictures posted by the user and viewable by the user's contacts (e.g., "friends"). Alternatively, public galleries may be created by an administrator of the social messaging application 910, which includes media from any user of the application (and accessible to all users). In yet another embodiment, the social messaging application 910 may include a "magazine" feature, which includes articles and other content generated by publishers on the social messaging application's platform and accessible to any user. Any of these environments or platforms can be used to implement the concepts of embodiments of this disclosure.

[0095] In some embodiments, the short message system may include a message containing a short video clip or image that is deleted after a deletion trigger event, such as viewing time or viewing completion. In this embodiment, when the short video clip is acquired by device 900, the device implementing the facial modeling system 160 may identify, track, extract, and generate a facial representation within the short video clip, and send the short video clip to another device using the short message system.

[0096] Software Architecture

[0097] Figure 10 This is a block diagram 1000 showing the architecture of software 1002 that can be installed on the aforementioned device. Figure 10This is merely a non-limiting example of a software architecture, and it will be understood that many other architectures can be implemented to facilitate the functionality described herein. In various embodiments, software 1002 is composed of, for example... Figure 11 The hardware implementation of machine 1100 includes processor 1110, memory 1130, and I / O components 1150. In this example architecture, software 1002 can be conceptualized as a stack of layers, each providing specific functionality. For example, software 1002 includes layers such as operating system 1004, libraries 1006, frameworks 1008, and application 1010. Operationally, consistent with some embodiments, application 1010 invokes application programming interface (API) calls 1012 through the software stack and receives messages 1014 in response to API call 1012.

[0098] In various implementations, operating system 1004 manages hardware resources and provides public services. Operating system 1004 includes, for example, a kernel 1020, services 1022, and drivers 1024. Consistent with some embodiments, kernel 1020 serves as an abstraction layer between hardware and other software layers. For example, kernel 1020 provides functions such as memory management, processor management (e.g., scheduling), component management, network connectivity, and security settings. Services 1022 can provide other public services to other software layers. According to some embodiments, driver 1024 is responsible for controlling the underlying hardware or interfaced with the underlying hardware. For example, driver 1024 may include a display driver, camera driver, Bluetooth® driver, flash memory driver, serial communication driver (e.g., Universal Serial Bus (USB) driver), Wi-Fi® driver, audio driver, power management driver, etc.

[0099] In some embodiments, library 1006 provides low-level general-purpose infrastructure utilized by application 1010. Library 1006 may include system library 1030 (e.g., the C standard library), which provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Furthermore, library 1006 may include API libraries 1032, such as media libraries (e.g., libraries supporting the rendering and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codecs, Joint Picture Experts Group (JPEG or JPG), or Portable Web Graphics (PNG)), graphics libraries (e.g., OpenGL frameworks for rendering two-dimensional (2D) and three-dimensional (3D) graphics content on a display), database libraries (e.g., SQLite, which provides various relational database functionalities), web libraries (e.g., WebKit, which provides web browsing functionality), etc. Library 1006 may also include a wide variety of other libraries 1034 to provide application 1010 with many other APIs.

[0100] According to some embodiments, framework 1008 provides a high-level common architecture that can be utilized by application 1010. For example, framework 1008 provides various graphical user interface (GUI) functions, high-level resource management, advanced location nodes, etc. Framework 1008 can provide a wide range of other APIs that can be utilized by application 1010, some of which may be specific to a particular operating system or platform.

[0101] In example embodiments, application 1010 includes home application 1050, contact application 1052, browser application 1054, book reader application 1056, location application 1058, media application 1060, messaging application 1062, game application 1064, and other broadly categorized applications such as third-party application 1066. According to some embodiments, application 1010 is a program that performs functions defined in a program. Application 1010 can be created using various programming languages, such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language). In specific examples, third-party application 1066 (e.g., an application developed by an entity other than a platform-specific vendor using the Android™ or iOS™ Software Development Kit (SDK)) can be mobile software that runs on a mobile operating system such as iOS™, Android™, Windows©Phone, or other mobile operating systems. In this example, a third-party application 1066 can call API call 1012 provided by the operating system 1004 to perform the functions described herein.

[0102] Example machine architecture and machine-readable media

[0103] Figure 11 This is a block diagram illustrating components of a machine 1100, according to some embodiments, capable of reading instructions (e.g., processor-executable instructions) from a machine-readable storage medium (e.g., a machine-readable storage medium) and performing any of the methods discussed herein. Specifically, Figure 11 A schematic diagram of machine 1100 in the form of an example computer system is shown, within which instructions 1116 (e.g., software, programs, applications, applets, or other executable code) can be executed to cause machine 1100 to perform any of the methods discussed herein. In alternative embodiments, machine 1100 operates as a standalone device or can be coupled (e.g., network-connected) to other machines. In a networked deployment, machine 1100 can operate as a server machine or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1100 can include, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, network devices, network routers, network switches, network bridges, or any machine capable of executing instructions 1116 that continuously or otherwise specifies the actions that machine 1100 will take. Furthermore, although only a single machine 1100 is shown, the term "machine" can also be considered to include a collection of machines 1100 that individually or in combination execute instructions 1116 to perform any of the methods discussed herein.

[0104] In various embodiments, machine 1100 includes processor 1110, memory 1130, and I / O components 1150 configured to communicate with each other via bus 1102. In example embodiments, processor 1110 (e.g., a central processing unit (CPU), a simplified instruction set computing (RISC) processor, a composite instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) includes, for example, processors 1112 and 1114 capable of executing instructions 1116. The term "processor" is intended to include multi-core processors that may include two or more independent processors (also referred to as "cores") capable of executing instructions 1116 simultaneously. Although Figure 11Multiple processors 1110 are shown, but machine 1100 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with single cores, multiple processors with multiple cores, or any combination thereof.

[0105] According to some embodiments, memory 1130 includes main memory 1132, static memory 1134, and storage cells 1136 accessible by processor 1110 via bus 1102. Storage cells 1136 may include machine-readable storage medium 1138 on which instructions 1116 embodying any of the methods or functions described herein are stored. Instructions 1116 may also reside wholly or at least partially in main memory 1132, static memory 1134, at least one of processor 1110 (e.g., in the processor's cache memory), or any suitable combination thereof during execution by machine 1100. Therefore, in various embodiments, main memory 1132, static memory 1134, and processor 1110 are considered to be machine-readable medium 1138.

[0106] As used herein, the term "memory" refers to machine-readable storage medium 1138 capable of temporarily or permanently storing data, and may be considered to include, but is not limited to, random access memory (RAM), read-only memory (ROM), cache, flash memory, and cache. While machine-readable storage medium 1138 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media capable of storing instructions 1116 (e.g., a centralized or distributed database, or associated caches and servers). The term "machine-readable storage medium" can also be considered to include any medium or combination of media capable of storing instructions (e.g., instructions 1116) for execution by a machine (e.g., machine 1100), such that when executed by a processor (e.g., processor 1110) of machine 1100, the instructions cause machine 1100 to perform any of the methods described herein. Therefore, "machine-readable storage medium" refers to a single storage device or apparatus, as well as a "cloud-based" storage system or storage network that includes multiple storage devices or apparatuses. Therefore, the term "machine-readable storage medium" can be considered as, but is not limited to, a data repository in the form of solid-state memory (e.g., flash memory), optical media, magnetic media, other non-volatile memory (e.g., erasable programmable read-only memory (EPROM)), or any suitable combination thereof. The term "machine-readable storage medium" specifically excludes non-legal signals themselves.

[0107] I / O component 1150 includes a variety of components for receiving input, providing output, generating output, sending information, exchanging information, acquiring measurements, etc. Generally, it is understood that I / O component 1150 may include... Figure 11 Many other components are not shown. I / O components 1150 are grouped according to function only for the purpose of simplifying the following discussion, and the grouping is by no means limiting. In various example embodiments, I / O components 1150 include output components 1152 and input components 1154. Output components 1152 include visual components (e.g., displays, such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), auditory components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. Input components 1154 include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other indicating instruments), tactile input components (e.g., physical buttons, touchscreens that provide the position and force of a touch or touch gesture, or other tactile input components), audio input components (e.g., microphones), etc.

[0108] In some other example embodiments, I / O component 1150 includes a variety of other components such as biometric component 1156, motion component 1158, environmental component 1160, or position component 1162. For example, biometric component 1156 includes components for detecting expressions (e.g., hand gestures, facial expressions, vocal expressions, body posture, or mouth gestures), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 1158 includes accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Environmental component 1160 includes, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., a thermometer that detects ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., a microphone that detects background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor component (e.g., a machine olfactory detection sensor, a gas detection sensor for detecting the concentration of hazardous gases or measuring pollutants in the atmosphere for safety purposes), or other components that may provide indications, measurements, or signals corresponding to the surrounding physical environment. Position component 1162 includes a positioning sensor component (e.g., a Global Positioning System (GPS) receiver component), an altitude sensor component (e.g., an altimeter or barometer that can detect from which altitude the air pressure is derived), an orientation sensor component (e.g., a magnetometer), etc.

[0109] Communication can be implemented using a wide variety of technologies. I / O component 1150 may include communication component 1164, operable to couple machine 1100 to network 1180 or device 1170 via coupler 1182 and coupler 1172, respectively. For example, communication component 1164 includes a network interface component or another suitable device interfaced with network 1180. In further examples, communication component 1164 includes wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth® components (e.g., Bluetooth Low Energy®), Wi-Fi® components, and other communication components that provide communication via other modes. Device 1170 may be another machine or any of a wide variety of peripheral devices, such as peripheral devices coupled via Universal Serial Bus (USB).

[0110] Furthermore, in some embodiments, the communication component 1164 detects identifiers or includes components operable to detect identifiers. For example, the communication component 1164 includes a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, data matrices, digital graphics, maximum codes, PDF417, supercodes, Uniform Business Code Reduced Space Symbol (UCC RSS)-2D barcodes, and other optical codes), an acoustic detection component (e.g., a microphone for identifying audio signals from the tag), or any suitable combination thereof. Additionally, various information can be derived via the communication component 1164, which can indicate a specific location, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location via detecting Bluetooth® or NFC beacon signals, etc.

[0111] transmission medium

[0112] In various example embodiments, a portion of network 1180 may be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a Common Old-Style Telephone Service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi® network, another type of network, or a combination of two or more such networks. For example, network 1180 or a portion of network 1180 may include a wireless or cellular network, and coupling 1182 may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or another type of cellular or wireless coupling. In this example, coupling 1182 can implement any of a variety of data transmission technologies, such as single-carrier radio transmission technology (1xRTT), evolved data optimization (EVDO) technology, general packet radio service (GPRS) technology, GSM evolution enhanced data rate (EDGE) technology, the 3rd generation partnership program (3GPP) including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high-speed packet access (HSPA), global microwave access interoperability (WiMAX), long-term evolution (LTE) standard, other standards defined by various standards-setting organizations, other remote protocols or other data transmission technologies.

[0113] In an example embodiment, instructions 1116 are sent or received over network 1180 via a transmission medium using a network interface device (e.g., a network interface component included in communication component 1164), and utilizing any of a plurality of known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, in other example embodiments, instructions 1116 are sent or received to device 1170 via a transmission medium using a coupling 1172 (e.g., peer-to-peer coupling). The term “transmission medium” can be considered to include any intangible medium capable of storing, encoding, or carrying instructions 1116 executed by machine 1100, and includes digital or analog communication signals or other intangible media to facilitate the communication implementation of such software.

[0114] Furthermore, because the machine-readable storage medium 1138 does not embody a propagating signal, it can be non-transient (in other words, it does not have any short-term signals). However, labeling the machine-readable storage medium 1138 as "non-transient" should not be interpreted as meaning that the medium cannot be moved. The medium should be considered as transferable from one physical location to another. In some embodiments, the machine-readable storage medium includes machine-readable stored signals. Additionally, since the machine-readable storage medium 1138 is tangible, it can be considered a machine-readable device. A transmission medium is an embodiment of a machine-readable medium.

[0115] The examples numbered below are implementation examples.

[0116] 1. A method comprising: receiving, by one or more processors, one or more images depicting at least a portion of one or more faces; detecting, by the one or more processors, a portion of the one or more faces depicted within the one or more images; in response to detecting each portion of the one or more faces, identifying a set of facial features depicted on the portion of the one or more faces depicted within the one or more images; generating, based on the identified set of facial features, a descriptor function representing the set of facial features, the descriptor function including a set of object functions, each object function representing a facial feature of the set of facial features; in response to generating the descriptor function, fitting each object function of the set of object functions; identifying a recognition probability for each facial feature of the set of facial features, the recognition probability being the probability that the facial feature corresponds to a specified characteristic in a set of feature characteristics; and assigning an identifier to each facial feature based on the recognition probability identified for each facial feature.

[0117] 2. The method according to Example 1 further includes one or more processors accessing a reference image from a facial reference database, the reference image having a set of reference facial landmarks; and in response to identifying the set of facial features of each of the one or more faces, mapping each face and the set of facial features of each face to the reference image.

[0118] 3. The method according to Example 1 or 2, wherein the reference image is a synthetic face representing a standardized face generated from multiple faces.

[0119] 4. The method according to any one or more of Examples 1-3, wherein fitting each object function in a set of object functions to the descriptor function further includes: updating one or more object functions within the descriptor function using stochastic gradient descent.

[0120] 5. The method according to any one or more of Examples 1-4, wherein the set of object functions comprises a plurality of object functions, and modifying one or more object functions further comprises selecting a first object function as a regularization function; and modifying one or more remaining object functions of the plurality of object functions relative to the regularization function.

[0121] 6. The method according to any one or more of Examples 1-5, wherein the recognition probability is a recognition probability value defined between a first value and a second value, the first value representing male and the second value representing female, and the method further includes recognizing a gender threshold between the first value and the second value; and selecting the gender for a face based on the recognition probability value moving toward either the first value or the second value exceeding the gender threshold.

[0122] 7. The method according to any one or more of Examples 1-6, wherein a first designated object function in a set of object functions is a convolutional layer describing at least one facial feature in a set of facial features, and the method further includes: dividing one or more images into a set of pixel regions; performing a convolutional layer for each pixel region in the set of pixel regions by performing the first designated object function on each pixel region; and generating a value representing the presence of a visual feature within the pixel region for each pixel region in response to performing the convolutional layer on each pixel region.

[0123] 8. The method according to any one or more of Examples 1-7, wherein one or more images are represented as a two-dimensional matrix, and the method further includes generating a resulting matrix including values ​​representing the presence of visual features within pixel regions.

[0124] 9. The method according to any one or more of Examples 1-8, wherein a second specified object function of a set of object functions is a pooling layer configured to reduce the size of the resulting matrix, and further includes: identifying the maximum value of one or more sub-regions of each pixel region; selecting the maximum value of each sub-region as a representation of the pixel region including the sub-region in the resulting matrix; and generating a pooling matrix including the maximum value of each pixel region.

[0125] 10. The method according to any one or more of Examples 1-9, wherein a third specified object function in a set of object functions is a connection layer representing the connection between the result matrix and the pooling matrix, and further includes: generating an inner product of the pooling matrix and the layer parameters; and generating a recognition probability based on the inner product, the pooling matrix, and the result matrix.

[0126] 11. A system comprising: one or more processors; and a processor-readable storage device coupled to the one or more processors and storing processor-executable instructions, which, when executed by the one or more processors, cause the one or more processors to perform operations including: receiving by the one or more processors one or more images depicting at least a portion of one or more faces; detecting by the one or more processors portions of the one or more faces depicted within the one or more images; identifying a set of facial features depicted on the portions of the one or more faces depicted within the one or more images in response to detecting each portion of the one or more faces; generating a descriptor function representing the set of facial features based on the identified set of facial features, the descriptor function including a set of object functions, each object function representing a facial feature of the set of facial features; fitting each object function of the set of object functions in response to generating the descriptor function; identifying a recognition probability of each facial feature of the set of facial features, the recognition probability being the probability that the facial feature corresponds to a specified characteristic in a set of feature characteristics; and assigning an identifier to each facial feature based on the recognition probability for each facial feature.

[0127] 12. The system according to Example 11, wherein fitting each of the set of object functions to the descriptor function further includes: updating one or more object functions within the descriptor function using stochastic gradient descent.

[0128] 13. The system according to Example 11 or 12, wherein a set of object functions comprises a plurality of object functions, and modifying one or more object functions further comprises: selecting a first object function as a regularization function; and modifying one or more of the remaining object functions among the plurality of object functions relative to the regularization function.

[0129] 14. A system according to any one or more of Examples 11-13, wherein a first designated object function in a set of object functions is a convolutional layer describing at least one facial feature of a set of facial features, and the operation further includes: dividing one or more images into a set of pixel regions; performing a convolutional layer for each pixel region of the set of pixel regions by performing the first designated object function on each pixel region; generating values ​​representing the presence of visual features within the pixel region for each pixel region in response to performing the convolutional layer on each pixel region; and generating a result matrix including values ​​representing the presence of visual features within the pixel regions based on one or more images represented as a two-dimensional matrix.

[0130] 15. A system according to any one or more of Examples 11-14, wherein a second designated object function of a set of object functions is a pooling layer configured to reduce the size of the resulting matrix, and the operation further includes: identifying the maximum value of one or more sub-regions of each pixel region; selecting the maximum value of each sub-region as a representation of the pixel region including the sub-region in the resulting matrix; and generating a pooling matrix including the maximum value of each pixel region.

[0131] 16. A system according to any one or more of Examples 11-15, wherein a third specified object function of a set of object functions is a connection layer representing the connection between the result matrix and the pooling matrix, and further includes: generating an inner product of the pooling matrix and the layer parameters; and generating a recognition probability based on the inner product, the pooling matrix, and the result matrix.

[0132] 17. A machine-readable storage medium carrying processor-executable instructions that, when executed by a processor of a machine, cause the machine to perform operations including: receiving by one or more processors one or more images depicting at least a portion of one or more faces; detecting by one or more processors portions of one or more faces depicted within the one or more images; identifying a set of facial features depicted on the portions of the faces depicted within the one or more images in response to detecting each portion of the one or more faces; generating a descriptor function representing the set of facial features based on the identified set of facial features, the descriptor function including a set of object functions, each object function representing a facial feature in the set of facial features; fitting each object function in the set of object functions in response to generating the descriptor function; identifying a recognition probability of each facial feature in the set of facial features, the recognition probability being the probability that the facial feature corresponds to a specified characteristic of a set of feature characteristics; and assigning an identifier to each facial feature based on the recognition probability identified for each facial feature.

[0133] 18. The machine-readable storage medium according to Example 17, wherein a first designated object function in a set of object function filters is a convolutional layer describing at least one facial feature in a set of facial features, and the operation further includes: dividing one or more images into a set of pixel regions; performing a convolutional layer for each pixel region in the set of pixel regions by performing the first designated object function on each pixel region; generating values ​​representing the presence of visual features within the pixel region for each pixel region in response to performing the convolutional layer on each pixel region; and generating a result matrix including values ​​representing the presence of visual features within the pixel regions based on one or more images represented as a two-dimensional matrix.

[0134] 19. The machine-readable storage medium according to Example 17 or 18, wherein a second designated object function of a set of object functions is a pooling layer configured to reduce the size of the resulting matrix, and the operation further includes: identifying the maximum value of one or more sub-regions of each pixel region; selecting the maximum value of each sub-region as a representation of the pixel region including the sub-region in the resulting matrix; and generating a pooling matrix including the maximum value of each pixel region.

[0135] 20. A machine-readable storage medium according to any one or more of Examples 17-19, wherein a third specified object function of a set of object functions is a connection layer representing a connection between a result matrix and a pooling matrix, and the operation further includes: generating an inner product of the pooling matrix and layer parameters; and generating a recognition probability based on the inner product, the pooling matrix, and the result matrix.

[0136] 21. A machine-readable medium carrying instructions that, when executed by one or more processors of a machine, cause the machine to perform the method according to any one of Examples 1 to 10.

[0137] language

[0138] Throughout this specification, multiple instances can implement components, operations, or structures described as single instances. While individual operations of methods are shown and described as separate operations, these individual operations can be performed simultaneously, and they do not need to be performed in the order shown. Structures and functionalities presented as individual components in the example configuration can be implemented as combined structures or components. Similarly, structures and functionalities presented as single components can be implemented as multiple separate components. These and other variations, modifications, additions, and improvements fall within the scope of this document's subject matter.

[0139] While an overview of the subject matter of the invention has been described with reference to specific exemplary embodiments, various modifications and changes may be made to these embodiments without departing from the broader scope of the embodiments of this disclosure. Such embodiments of the subject matter of the invention may be referred to herein individually or collectively by the term "invention," which is merely for convenience and is not intended to limit the scope of this application to any single disclosure or inventive concept if more than one is disclosed in fact.

[0140] The embodiments shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments may be used and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Therefore, the specific implementation should not be considered limiting, and the scope of the various embodiments is defined only by the appended claims and the full scope of their equivalents.

[0141] As used herein, the term "or" may be interpreted in an inclusive or exclusive manner. Furthermore, multiple instances of the resources, operations, or structures described herein may be provided as a single instance. Moreover, the boundaries between various resources, operations, components, engines, and data stores are somewhat arbitrary, and specific operations are illustrated in the context of a particular illustrative configuration. Other allocations of functionality are contemplated, and these other allocations may fall within the scope of various embodiments of this disclosure. Typically, structures and functions presented as separate resources in example configurations may be implemented as combined structures or resources. Similarly, structures and functions presented as single resources may be implemented as separate resources. These and other variations, modifications, additions, and improvements fall within the scope of embodiments of this disclosure as represented by the appended claims. Therefore, the specification and drawings are to be considered illustrative rather than restrictive.

Claims

1. A method for facial modeling, comprising: Features of objects depicted within an image are identified by one or more processors, the objects including faces and the features including a set of facial features; A descriptor function representing the set of facial features is generated. The descriptor function includes multiple object functions, and each object function represents a different facial feature in the set of facial features. The descriptor function and the multiple object functions represent a neural network structure, and each layer of the neural network structure represents an object function. In response to generating the descriptor function, each of the plurality of object functions is fitted relative to each other; Obtain a probability boundary that includes a first value and a second value corresponding to two or more feature values ​​corresponding to a first recognition aspect and a second recognition aspect of facial features in the set of facial features of a given object, wherein the probability boundary has a desired minimum boundary value and a maximum boundary value; The object function is used to calculate the numerical probability of a facial feature in the set of facial features of the object depicted in the image, the numerical probability indicating the probability that the facial feature corresponds to a specified recognition aspect in a set of recognition aspects of the facial features of the given object; Determine whether the numerical probability of the facial feature of the object is closer in proximity to a first or second value included in the probability boundary, to indicate whether the facial feature corresponds to a first or second recognition aspect of the facial feature of the given object; and Based on the determination of whether the numerical probability of the facial features of the object depicted in the image is closer in proximity to a first or second value included in the probability boundary, an identifier is assigned to the facial features of the object.

2. The method according to claim 1, further comprising: The one or more processors access reference images from an object reference database, the reference images having a set of reference object landmarks; as well as In response to recognizing the facial features, the object and the facial features of the object are mapped onto the reference image.

3. The method according to claim 1, wherein, Fitting the object function includes using stochastic gradient descent updates to modify one or more of the object functions within the descriptor function.

4. The method according to claim 3, wherein, Modifying one or more of the object functions further includes: Select the first object function as the regularization function; and The regularization function modifies one or more of the remaining object functions among the plurality of object functions.

5. The method according to claim 1, wherein, The first value represents male, and the second value represents female, the method further comprising: Identify a gender threshold between the first value and the second value; and The gender is selected based on the recognition probability value exceeding the gender threshold, moving towards either the first value or the second value.

6. The method of claim 1, further comprising: Provides a convolutional layer that describes at least one facial feature from a set of facial features; The image is divided into a set of pixel regions; The convolutional layer is executed for each pixel region in the set of pixel regions by executing a specified object function on each pixel region; as well as In response to performing the convolutional layer on each pixel region, a value representing the presence of visual features within the pixel region is generated for each pixel region.

7. The method according to claim 6, wherein, The image is represented as a two-dimensional matrix and further includes: Generate a result matrix that includes values ​​representing the presence of the visual features within the pixel region.

8. The method according to claim 7, wherein, The specified object function is a pooling layer configured to reduce the size of the resulting matrix, and further includes: Identify the maximum value of one or more sub-regions within each pixel region; Select the maximum value of each sub-region as a representation of the pixel region including the sub-region in the resulting matrix; and Generate a pooling matrix that includes the maximum value for each pixel region.

9. The method according to claim 8, wherein, Another specified object function is a connection layer representing the connection between the resulting matrix and the pooling matrix, and further includes: Generate the inner product of the pooling matrix and the layer parameters; and The numerical probability is generated based on the inner product, the pooling matrix, and the result matrix.

10. A system for facial modeling, comprising: One or more processors; as well as A processor-readable storage device coupled to the one or more processors and storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including the following: Features of objects depicted within an image are identified by one or more processors, the objects including faces and the features including a set of facial features; A descriptor function representing the set of facial features is generated. The descriptor function includes multiple object functions, and each object function represents a different facial feature in the set of facial features. The descriptor function and the multiple object functions represent a neural network structure, and each layer of the neural network structure represents an object function. In response to generating the descriptor function, each of the plurality of object functions is fitted relative to each other; Obtain a probability boundary that includes a first value and a second value corresponding to two or more feature values ​​corresponding to a first recognition aspect and a second recognition aspect of facial features in the set of facial features of a given object, wherein the probability boundary has a desired minimum boundary value and a maximum boundary value; The object function is used to calculate the numerical probability of a facial feature in the set of facial features of the object depicted in the image, the numerical probability indicating the probability that the facial feature corresponds to a specified recognition aspect in a set of recognition aspects of the facial features of the given object; Determine whether the numerical probability of the facial feature of the object is closer in proximity to a first or second value included in the probability boundary, to indicate whether the facial feature corresponds to a first or second recognition aspect of the facial feature of the given object; and Based on the determination of whether the numerical probability of the facial features of the object depicted in the image is closer in proximity to a first or second value included in the probability boundary, an identifier is assigned to the facial features of the object.

11. The system according to claim 10, wherein: Fitting each of the object functions in the set of object functions to the descriptor function further includes using stochastic gradient descent to modify one or more of the object functions within the descriptor function; The operation further includes updating the functions of the group of objects as follows: Select the first object function as the regularization function; and The regularization function modifies one or more of the remaining object functions among the plurality of object functions.

12. The system according to claim 10, wherein, The operation further includes: Provides a convolutional layer that describes at least one facial feature from a set of facial features; The image is divided into a set of pixel regions; The convolutional layer is executed for each pixel region in the set of pixel regions by executing a specified object function on each pixel region; In response to performing the convolutional layer on each pixel region, a value representing the presence of visual features within the pixel region is generated for each pixel region; and Based on the image, which is represented as a two-dimensional matrix, a result matrix is ​​generated including values ​​representing the presence of the visual features within the pixel region.

13. The system according to claim 12, wherein, The specified object function is a pooling layer configured to reduce the size of the resulting matrix, and the operation further includes: Identify the maximum value of one or more sub-regions within each pixel region; Select the maximum value of each sub-region as a representation of the pixel region including the sub-region in the resulting matrix; and Generate a pooling matrix that includes the maximum value for each pixel region.

14. The system according to claim 13, wherein, Another specified object function is a connection layer representing the connection between the resulting matrix and the pooling matrix, and further includes: Generate the inner product of the pooling matrix and the layer parameters; and The numerical probability is generated based on the inner product, the pooling matrix, and the result matrix.

15. A processor-readable storage device storing processor-executable instructions, which, when executed by a processor of a machine, cause the machine to perform operations including: Features of objects depicted within an image are identified by one or more processors, the objects including faces and the features including a set of facial features; A descriptor function representing the set of facial features is generated. The descriptor function includes multiple object functions, and each object function represents a different facial feature in the set of facial features. The descriptor function and the multiple object functions represent a neural network structure, and each layer of the neural network structure represents an object function. In response to generating the descriptor function, each of the plurality of object functions is fitted relative to each other; Obtain a probability boundary that includes a first value and a second value corresponding to two or more feature values ​​corresponding to a first recognition aspect and a second recognition aspect of facial features in the set of facial features of a given object, wherein the probability boundary has a desired minimum boundary value and a maximum boundary value; The object function is used to calculate the numerical probability of a facial feature in the set of facial features of the object depicted in the image, the numerical probability indicating the probability that the facial feature corresponds to a specified recognition aspect in a set of recognition aspects of the facial features of the given object; Determine whether the numerical probability of the facial feature of the object is closer in proximity to a first or second value included in the probability boundary, to indicate whether the facial feature corresponds to a first or second recognition aspect of the facial feature of the given object; and Based on the determination of whether the numerical probability of the facial features of the object depicted in the image is closer in proximity to a first or second value included in the probability boundary, an identifier is assigned to the facial features of the object.

16. The processor-readable storage device according to claim 15, wherein, The operation further includes: Provides a convolutional layer that describes at least one facial feature from a set of facial features; The image is divided into a set of pixel regions; The convolutional layer is executed for each pixel region in the set of pixel regions by executing a specified object function on each pixel region; and In response to performing the convolutional layer on each pixel region, a value representing the presence of visual features within the pixel region is generated for each pixel region; and Generate a result matrix based on the one or more images represented as two-dimensional matrices, including values ​​representing the presence of the visual features within the pixel regions.

17. The processor-readable storage device according to claim 16, wherein, The specified object function is a pooling layer configured to reduce the size of the resulting matrix, and the operation further includes: Identify the maximum value of one or more sub-regions within each pixel region; Select the maximum value of each sub-region as a representation of the pixel region including the sub-region in the resulting matrix; and Generate a pooling matrix that includes the maximum value for each pixel region.

18. The processor-readable storage device according to claim 17, wherein, Another specified object function is a connection layer representing the connection between the resulting matrix and the pooling matrix, and the operation further includes: Generate the inner product of the pooling matrix and the layer parameters; and The numerical probability is generated based on the inner product, the pooling matrix, and the result matrix.

19. A method for facial modeling, comprising: An image is segmented into a set of pixel regions by one or more processors, the set comprising multiple pixel regions, wherein a given object is depicted within the image, the given object including a face and the features of the given object including a set of facial features; A descriptor function is generated to represent a set of facial features of the given object. The descriptor function includes multiple object functions, and each object function represents a different facial feature in the set of facial features. The descriptor function and the multiple object functions represent a neural network structure, and each layer of the neural network structure represents an object function. In response to generating the descriptor function, each of the plurality of object functions is fitted relative to each other; The first object function of the plurality of object functions is executed by the one or more processors on the set of pixel regions; The one or more processors generate values ​​representing the presence of visual features within a pixel region of the set of pixel regions. These values ​​representing the presence of visual features are generated as a result of executing the first object function on the set of pixel regions. The values ​​representing the presence of visual features are also generated based on a specified measurement of the proximity of the visual feature to one of a first value and a second value included in a probability boundary. The first and second values ​​included in the probability boundary correspond to two or more feature values ​​that correspond to a recognition aspect of the visual feature of the given object, including facial features. The one or more processors use the object function to generate a result matrix including values ​​representing the presence of the visual features within the pixel region; Identify the facial features of the given object depicted within the image; Obtain the probability boundary that includes the first value and the second value, wherein the probability boundary has a desired minimum boundary value and a maximum boundary value; Based on the result matrix, calculate the numerical probability of a facial feature in the set of facial features of the given object depicted in the image, the numerical probability indicating the probability that the facial feature corresponds to a specified recognition aspect in a set of recognition aspects of the visual features of the given object; Determine whether the numerical probability of the facial feature of the given object is closer in proximity to a first or second value included in the probability boundary, to indicate whether the facial feature corresponds to a first or second recognition aspect of the visual feature of the given object; and Based on the determination of whether the numerical probability of the facial features of the given object depicted in the image is closer in proximity to a first or second value included in the probability boundary, an identifier is assigned to the facial features of the given object.

20. The method according to claim 19, wherein, Executing the first object function includes: Identify facial features of the given object within a portion of the set of pixel regions; and The first object function is selected from a plurality of object functions based on the facial features, wherein each of the plurality of object functions corresponds to a different facial feature.

21. The method of claim 19, further comprising: The one or more processors access reference images from an object reference database, the reference images having a set of reference object landmarks; as well as Map the objects depicted in the image to the reference image.

22. The method of claim 19, further comprising fitting a plurality of object functions in a cascaded manner, wherein each fitted object function causes a modification to the parameters of one or more subsequent object functions, each of the plurality of object functions identifying different facial features of an object described in the image.

23. The method of claim 22, wherein the first object function of the plurality of object functions identifies the race of a face in the image, wherein the second object function of the plurality of object functions identifies the gender of the face in the image, and wherein the third object function of the plurality of object functions identifies the age of the face in the image.

24. The method according to claim 19, wherein, The first value represents male, and the second value represents female, the method further comprising: Identify a gender threshold between the first value and the second value; and The sex is selected based on the probability value exceeding the sex threshold, moving towards either the first or the second value.

25. The method according to claim 23, wherein, Another specified object function is a connection layer representing the connection between the resulting matrix and the pooling matrix, and further includes: Generate the inner product of the pooling matrix and the layer parameters; and Numerical probabilities are generated based on the inner product, the pooling matrix, and the result matrix.

26. The method according to claim 19, wherein, The object function includes a convolutional layer as a filter operation, and further includes: The size of the filter is determined based on the dimensions or measurements of the image; and The filter is used to process each pixel in the image, the processing including: Identify one or more values ​​of the pixel; Extract the region surrounding the pixel whose size is determined according to the filter; and A single output is generated for the pixel based on the values ​​of the extracted region.

27. The method according to claim 19, wherein, The generated value indicates the presence of the visual feature, wherein the determined value of the pixel region exceeds a threshold.

28. The method of claim 19, further comprising: Identify the maximum value of one or more sub-regions within each pixel region of this group of pixel regions; Select the maximum value of each of the one or more sub-regions as the representation of the pixel region including the sub-region; as well as Generate a pooling matrix that includes the maximum value of each pixel region in the set of pixel regions.

29. The method according to claim 28, wherein, The maximum value is identified by executing a second object function configured to reduce the size of the resulting matrix.

30. A system for facial modeling, comprising: One or more processors; as well as A processor-readable storage device coupled to the one or more processors and storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including the following: The image is segmented into a set of pixel regions, which includes multiple pixel regions, wherein a given object is depicted within the image, the given object including a face and the features of the given object including a set of facial features; A descriptor function is generated to represent a set of facial features of the given object. The descriptor function includes multiple object functions, and each object function represents a different facial feature in the set of facial features. The descriptor function and the multiple object functions represent a neural network structure, and each layer of the neural network structure represents an object function. In response to generating the descriptor function, each of the plurality of object functions is fitted relative to each other; Execute the first object function of the plurality of object functions on this set of pixel regions; A value representing the presence of a visual feature within a pixel region of the set of pixel regions is generated as a result of performing the first object function on the set of pixel regions. The value representing the presence of the visual feature is also generated based on a specified measurement of the proximity of the visual feature to one of a first value and a second value included in a probability boundary, the first value and the second value included in the probability boundary corresponding to two or more feature values ​​corresponding to the recognition aspect of the visual feature of the given object, the visual feature including facial features. The object function is used to generate a result matrix that includes values ​​representing the presence of the visual features within the pixel region; Identify the facial features of the given object depicted within the image; Obtain the probability boundary that includes the first value and the second value, wherein the probability boundary has a desired minimum boundary value and a maximum boundary value; Based on the result matrix, calculate the numerical probability of a facial feature in the set of facial features of the given object depicted in the image, the numerical probability indicating the probability that the facial feature corresponds to a specified recognition aspect in a set of recognition aspects of the visual features of the given object; Determine whether the numerical probability of the facial feature of the given object is closer in proximity to a first or second value included in the probability boundary, to indicate whether the facial feature corresponds to a first or second recognition aspect of the visual feature of the given object; and Based on the determination of whether the numerical probability of the facial features of the object depicted in the image is closer in proximity to a first or second value included in the probability boundary, an identifier is assigned to the facial features of the object.

31. The system according to claim 30, wherein, The specified object function represents a connection layer between the result matrix and the pooling matrix, further including: Generate the inner product of the pooling matrix and the layer parameters; and Numerical probabilities are generated based on the inner product, the pooling matrix, and the result matrix.

32. A processor-readable storage device storing processor-executable instructions, which, when executed by a processor of a machine, cause the machine to perform operations including: The image is segmented into a set of pixel regions, which includes multiple pixel regions, wherein a given object is depicted within the image, the given object including a face and the features of the given object including a set of facial features; A descriptor function is generated to represent a set of facial features of the given object. The descriptor function includes multiple object functions, and each object function represents a different facial feature in the set of facial features. The descriptor function and the multiple object functions represent a neural network structure, and each layer of the neural network structure represents an object function. In response to generating the descriptor function, each of the plurality of object functions is fitted relative to each other; Execute the first object function of the plurality of object functions on this set of pixel regions; A value representing the presence of a visual feature within a pixel region of the set of pixel regions is generated as a result of performing the first object function on the set of pixel regions. The value representing the presence of the visual feature is also generated based on a specified measurement of the proximity of the visual feature to one of a first value and a second value included in a probability boundary, the first value and the second value included in the probability boundary corresponding to two or more feature values ​​corresponding to the recognition aspect of the visual feature of the given object, the visual feature including facial features. The object function is used to generate a result matrix that includes values ​​indicating the presence of the facial features within the pixel region; Identify the facial features of the given object depicted within the image; Obtain the probability boundary that includes the first value and the second value, wherein the probability boundary has a desired minimum boundary value and a maximum boundary value; Based on the result matrix, calculate the numerical probability of a facial feature in the set of facial features of the given object depicted in the image, the numerical probability indicating the probability that the facial feature corresponds to a specified recognition aspect in a set of recognition aspects of the visual features of the given object; Determine whether the numerical probability of the facial feature of the given object is closer in proximity to a first or second value included in the probability boundary, to indicate whether the facial feature corresponds to a first or second recognition aspect of the visual feature of the given object; and Based on the determination of whether the numerical probability of the facial features of the object depicted in the image is closer in proximity to a first or second value included in the probability boundary, an identifier is assigned to the facial features of the object.

Citation Information

Patent Citations

  • Systems and methods for facial representation

    CN105874474A

  • Method and apparatus for recognizing object, and method and apparatus for learning recognizer

    KR1020160061856A

  • Method of detecting facial attributes

    WO2012139273A1