Input method and device
By collecting and identifying facial expressions, lip reading, gestures or sign language in user video data, and providing candidate options for expressing information, the problem of limited input in specific scenarios of existing input methods is solved, and a more flexible and personalized input method is achieved.
Patent Information
- Application Number
- CN201910828531.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-09-03
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2039-09-03
AI Technical Summary
Existing input methods cannot implement voice or pinyin input in some scenarios, resulting in limited user input.
By collecting user video data, identifying video data such as facial expressions, lip reading, gestures or sign language, the corresponding expression information is identified and displayed as candidate items for users to choose.
When voice or pinyin input is not possible, the input flexibility and user experience are improved to meet personalized needs.
Smart Images

Figure CN112446265B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, and in particular to an input method and device. Background Art
[0002] With the widespread adoption of terminal devices, users are increasingly performing a large number of input operations on these devices. This is typically accomplished using input methods, which are encoding methods used to enter various symbols into computers or other devices. In existing technologies, input methods can be used to achieve input through phonetic or voice input. However, these methods are limited by usage scenarios, such as when both hands are busy and speaking is inconvenient. Summary of the Invention
[0003] In view of this, embodiments of the present application provide an input method and device to achieve more flexible input.
[0004] To solve the above problems, the technical solutions provided in the embodiments of the present application are as follows:
[0005] In a first aspect of an embodiment of the present application, an input method is provided, wherein the method collects user video data, wherein the user video data includes user facial video data and / or user hand video data;
[0006] Identifying expression information included in the user video data;
[0007] The expression information is displayed as a candidate.
[0008] In a possible implementation, identifying the expression information included in the user video data includes:
[0009] Classifying the user video data to obtain a video classification result;
[0010] The expression information included in the user video data is identified according to the video classification result.
[0011] In a possible implementation, identifying the expression information included in the user video data according to the video classification result includes:
[0012] When the video classification result includes facial expression video data, determining corresponding expression information based on the user's facial features in the facial expression video data;
[0013] When the video classification result includes lip reading video data, determining pronunciation according to lip shape features in the lip reading video data, and determining corresponding expression information according to the pronunciation;
[0014] When the video classification result includes gesture video data, determining corresponding expression information according to user hand features included in the gesture video data;
[0015] When the video classification result includes sign language video data, corresponding semantics are determined according to user hand features in the sign language video data, and corresponding expression information is determined according to the semantics.
[0016] In a possible implementation, classifying the user video data to obtain a video classification result includes:
[0017] Extracting user facial features and / or user hand features from the user video data;
[0018] The video classification result is determined based on the user's facial features and / or the user's hand features.
[0019] In a possible implementation, determining the video classification result according to the user's facial features and / or the user's hand features includes:
[0020] When the user video data includes the user's facial features, if the lip shape features in the user's facial features continuously change, the video classification result includes lip reading video data; otherwise, the video classification result includes facial expression video data;
[0021] When the user video data includes the user hand feature, if the user hand feature changes continuously, the video classification result includes sign language video data; otherwise, the video classification result includes gesture video data.
[0022] In a possible implementation, identifying the expression information included in the user video data includes:
[0023] Identifying whether the user video data includes user facial features, and if so, determining corresponding expression information based on the user facial features;
[0024] Identifying whether the user video data includes user facial features, and if so, determining whether a lip shape feature in the user facial features continuously changes, and if so, determining a pronunciation based on the lip shape feature, and determining corresponding expression information based on the pronunciation;
[0025] Identifying whether the user video data includes a user hand feature, and if so, determining corresponding expression information based on the user hand feature;
[0026] Identify whether the user hand video data includes user hand features; if so, determine whether the user hand features continuously change; if so, determine corresponding semantics based on the user hand features, and determine corresponding expression information based on the semantics.
[0027] In a possible implementation, the expression information includes one or more of expression text, image, and expression.
[0028] In a possible implementation, the expression text includes a standard expression text and / or a custom expression text;
[0029] When the expression text includes a custom expression text, displaying the expression information as a candidate item includes:
[0030] The custom expression text is displayed preferentially among the candidate items.
[0031] In a second aspect of an embodiment of the present application, an input device is provided, the device comprising:
[0032] A collection unit, configured to collect user video data, wherein the user video data includes user facial video data and / or user hand video data;
[0033] an identification unit, configured to identify expression information included in the user video data;
[0034] A display unit is used to display the expression information as a candidate item.
[0035] In a possible implementation, the identification unit includes:
[0036] A classification subunit, configured to classify the user video data and obtain a video classification result;
[0037] The identification subunit is used to identify the expression information included in the user video data according to the video classification result.
[0038] In a possible implementation, the recognition subunit is specifically configured to, when the video classification result includes facial expression video data, determine corresponding expression information based on the user's facial features in the facial expression video data;
[0039] When the video classification result includes lip reading video data, determining pronunciation according to lip shape features in the lip reading video data, and determining corresponding expression information according to the pronunciation;
[0040] When the video classification result includes gesture video data, determining corresponding expression information according to user hand features included in the gesture video data;
[0041] When the video classification result includes sign language video data, corresponding semantics are determined according to user hand features in the sign language video data, and corresponding expression information is determined according to the semantics.
[0042] In a possible implementation, the classification subunit includes:
[0043] an extraction subunit, configured to extract user facial features and / or user hand features from the user video data;
[0044] A determination subunit is used to determine the video classification result based on the user's facial features and / or the user's hand features.
[0045] In a possible implementation, the determining subunit is specifically configured to, when the user video data includes the user facial features, if the lip shape features in the user facial features continuously change, the video classification result includes lip reading video data; otherwise, the video classification result includes facial expression video data;
[0046] When the user video data includes the user hand feature, if the user hand feature changes continuously, the video classification result includes sign language video data; otherwise, the video classification result includes gesture video data.
[0047] In a possible implementation, the identification unit includes:
[0048] a first recognition subunit, configured to identify whether the user video data includes user facial features, and if so, determine corresponding expression information based on the user facial features;
[0049] a second recognition subunit, configured to identify whether the user video data includes user facial features, and if so, determine whether a lip shape feature in the user facial features continuously changes; if so, determine a pronunciation based on the lip shape feature, and determine corresponding expression information based on the pronunciation;
[0050] a third identification subunit, configured to identify whether the user video data includes a user hand feature, and if so, determine corresponding expression information according to the user hand feature;
[0051] The fourth identification subunit is used to identify whether the user hand video data includes user hand features. If the user hand features are included, determine whether the user hand features change continuously. If the user hand features change continuously, determine the corresponding semantics based on the user hand features, and determine the corresponding expression information based on the semantics.
[0052] In a possible implementation, the expression information includes one or more of expression text, image, and expression.
[0053] In a possible implementation, the expression text includes a standard expression text and / or a custom expression text;
[0054] The display unit is specifically configured to, when the expression text includes a custom expression text, preferentially display the custom expression text among the candidate items.
[0055] In a third aspect of the embodiments of the present application, a device for input is provided, comprising a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the following operations:
[0056] Collecting user video data, where the user video data includes user facial video data and / or user hand video data;
[0057] Identifying expression information included in the user video data;
[0058] The expression information is displayed as a candidate.
[0059] In a fourth aspect of the embodiments of the present application, a computer-readable medium is provided, on which instructions are stored, which, when executed by one or more processors, enable the device to perform the input method as described in the first aspect.
[0060] It can be seen that the embodiments of the present application have the following beneficial effects:
[0061] In an embodiment of the present application, when it is inconvenient for the user to input voice or pinyin, the input method client collects user video data, which may include user facial video data and / or user hand video data. At the same time, the client identifies the user video data, obtains the expression information included in the user video data, and then displays the expression information as a candidate item, so that the user can select the candidate item to be expressed. That is, when it is inconvenient for the user to input pinyin or voice, the present application can complete the input of the sentence through facial video data, such as facial expressions, lip reading, etc., and / or user hand video data, such as gestures or sign language, etc., thereby improving the flexibility of input and the user's multi-faceted input experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 An example diagram of an application scenario provided in an embodiment of the present application;
[0063] Figure 2A flowchart of an input method provided in an embodiment of the present application;
[0064] Figure 3a An example diagram of a gesture provided in an embodiment of the present application;
[0065] Figure 3b An example diagram of facial expressions provided in an embodiment of the present application;
[0066] Figure 4 A structural diagram of an input device provided in an embodiment of the present application;
[0067] Figure 5 A structural diagram of another input device provided in an embodiment of the present application;
[0068] Figure 6 A server structure diagram provided for an embodiment of the present application. DETAILED DESCRIPTION
[0069] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0070] The inventors have found in their research on traditional input methods that current input methods mainly implement input using pinyin input or voice input, etc. However, in some application scenarios, voice or pinyin input cannot be used, which affects user input.
[0071] Based on this, an embodiment of the present application provides an input method, specifically, when the user is unable to use voice input and pinyin input, the input method client can collect user video data, and the user video data can include user facial video data and / or user hand video data. Then, the expression information included in the user video data is identified, and the expression information is displayed as a candidate item so that the user can select the required expression information from the candidate items. That is, when the user is not convenient to use voice or pinyin input, the user video data can be collected to identify the expression information included in the user video data through an image recognition method. In other words, the user can express through gestures, facial expressions, lip reading, etc. to achieve the input effect. In addition, in the present application, the user can customize the expression information corresponding to lip reading, gestures or facial expressions according to his or her actual situation to meet the user's personalized needs.
[0072] For ease of understanding of the embodiments of this application, see Figure 1 , which is a schematic diagram of a framework of an exemplary application scenario provided by an embodiment of the present application. The input method provided by an embodiment of the present application can be applied to an input method client 10.
[0073] In actual application, the client 10 collects user video data, identifies the user video data to obtain the expression information included in the user video data, and displays the expression information as a candidate item so that the user can select the expression information to be expressed and realize text input.
[0074] It should be noted that the recognition of the expression information included in the user video data can be performed by the client 10, or after the client 10 collects the user video data, it can send the user video data to the server 20, and the server 20 can recognize the expression information included in the user video data and send the recognition result to the client 10. The client 10 displays the acquired expression information as a candidate for the user to select the desired expression information.
[0075] The expression information may include one or more of expression text, images, and expressions, and the expression text may include standard expression text and / or custom expression text. In a specific implementation, when the expression text included in the user video data includes a custom expression text, the custom expression text is displayed first.
[0076] Those skilled in the art will understand that Figure 1 The framework diagram shown is only an example in which the embodiments of the present application can be implemented. The scope of application of the embodiments of the present application is not limited by any aspect of the framework.
[0077] It should be noted that the client 10 can be carried on a terminal, which can be any user device that is currently in development or will be developed in the future and can interact with each other through any form of wired and / or wireless connection (for example, Wi-Fi, LAN, cellular, coaxial cable, etc.), including but not limited to: smart wearable devices that are currently in development or will be developed in the future, smart phones, non-smart phones, tablet computers, laptop personal computers, desktop personal computers, minicomputers, mid-range computers, mainframe computers, etc. The implementation methods of the present application are not limited in this respect. It should also be noted that the server 20 in the embodiment of the present application can be an example of a device that is currently in development or will be developed in the future and can provide image recognition services. The implementation methods of the present application are not limited in this respect.
[0078] To facilitate understanding of the technical solution provided by the embodiment of the present application, the input method provided by the embodiment of the present application will be described below with reference to the accompanying drawings.
[0079] See also Figure 2 , which is a flow chart of an input method provided by an embodiment of the present application, such as Figure 2 As shown, the method includes:
[0080] S201: Collect user video data.
[0081] In this embodiment, when the user inputs through user video data, the client first collects the user video data, which may include user facial video data and / or user hand video data.
[0082] The user's facial video data may include lip reading video data and / or facial expression video data, i.e., the user may input through lip reading or facial expression. The user's hand video data may include sign language video data and / or gesture video data, i.e., the user may input through sign language or gesture.
[0083] In a specific implementation, the input method client may include an image acquisition function. When the user triggers this function, the input method client responds to the user's selection to start the image acquisition function and can complete the acquisition of the user's video data by calling the image acquisition device of the terminal.
[0084] S202: Identify expression information included in user video data.
[0085] In this embodiment, after the input method client collects user video data, it identifies the user video data to obtain the expression information included in the user video data. The expression information may include one or more of expression text, images, and expressions. That is, by identifying the user video data, one or more of the expression text, images, and expressions included in the user video data can be obtained. The expressions can be static or dynamic expression images.
[0086] It is understood that users can input through any one or more methods such as lip reading, facial expressions, gestures, sign language, etc. After the input method client obtains the user video data, it needs to identify the expression information corresponding to the collected user video data to obtain the expression information that the user may want to input. In specific implementations, the input method client can use its own image recognition function to identify the user video data to obtain the corresponding expression information, or it can send the collected user video data to the input method server, which will recognize it and then send the recognition result to the input method client.
[0087] In practical applications, user video data can be directly identified to obtain the corresponding expression information. Alternatively, the user video data can be first classified, and then the user video data can be identified based on the classification results to obtain the corresponding expression information, thereby improving the efficiency of identifying the expression information included in the user video data. The specific implementation of identifying the expression information included in the user video data will be described in subsequent embodiments.
[0088] S203: Display the expression information as a candidate.
[0089] In this embodiment, after the input method client obtains the expression information corresponding to the user video data, it displays the expression information as a candidate item. Since the expression information can include one or more of expression text, image, and expression, each candidate item can be any one of expression text, image, or expression.
[0090] It is understandable that when identifying the expression information included in the user video data, multiple expression information can be identified, and each expression information is displayed as a candidate to enable the user to select the candidate that best meets the user's expectations. For example, the user video data input by the user is the user's hand video data, such as Figure 3a As shown, the expression information included in the user hand video data obtained by identification includes expression text, which can be "OK", "No problem", "OK", etc., and the expression information included in the user hand video data obtained by identification also includes an expression image "including an expression of an OK gesture", then the text "OK", "No problem", "OK", and the expression image "including an expression of an OK gesture" are displayed as candidates.
[0091] Furthermore, to meet user needs, users can pre-customize the text corresponding to certain gestures, sign language, and facial expressions. Therefore, when identifying text included in user video data, the text can include standard text and / or customized text. If the text includes customized text, it will be displayed first among the candidate options, improving the user experience.
[0092] It can be seen from the above embodiments that when it is inconvenient for the user to input voice or pinyin, the client collects user video data, which may include user facial video data and / or user hand video data. At the same time, the client identifies the user video data, obtains the expression information included in the user video data, and then displays the expression information as a candidate item, so that the user can select the candidate item to be expressed. That is, when it is inconvenient for the user to input pinyin or voice, the present application can complete the input of sentences through facial video data, such as facial expressions, lip reading, etc., and / or user hand video data, such as gestures or sign language, thereby improving the flexibility of input and the user's multi-faceted input experience.
[0093] Based on the above description, it can be seen that the input method client can directly identify the expression information included in the user video data, or it can first classify the user video data and then classify the user video data based on the video classification results. For ease of understanding, the following will introduce each of them separately.
[0094] One method is to directly identify the user video data to obtain the expression information included in the user video data, specifically including:
[0095] 1) Identify whether the user video data includes the user's facial features. If the user's facial features are included, determine the corresponding expression information based on the user's facial features.
[0096] In this embodiment, after obtaining user video data, feature extraction is performed on the user video data, and it is identified whether the extracted features include user facial features. If user facial features are included, it indicates that the user may input through facial expressions, and the corresponding expression information can be determined based on the user facial features.
[0097] In specific implementation, an image recognition model that can recognize facial expressions (i.e., a facial expression image recognition model) can be pre-trained. In actual application, the facial expression image recognition model can be used to identify whether the extracted features include user facial features. If user facial features are included, the expression information included in the user video data can be identified.
[0098] 2) Identify whether the user video data includes the user's facial features. If the user's facial features are included, determine whether the lip shape features in the user's facial features change continuously. If the lip shape features change continuously, determine the pronunciation based on the lip shape features, and determine the corresponding expression information based on the pronunciation.
[0099] In this embodiment, after obtaining user video data, feature extraction is performed on the user video data, and it is identified whether the extracted features include user facial features. If the user facial features are included, it is determined whether the lip shape features in the user facial features are continuously changing. If the lip shape features are continuously changing, it indicates that the user may be inputting through lip reading. The pronunciation is determined based on the lip shape features, and the corresponding expression information is determined based on the pronunciation.
[0100] In specific implementation, an image recognition model that can recognize lip reading (i.e., a lip reading image recognition model) can be pre-trained. In actual application, the lip reading image recognition model can be used to identify whether the extracted features include user facial features. If the user facial features are included, it is determined whether the lip shape features in the user's facial features are continuously changing. If the lip shape features are continuously changing, the expression information included in the user video data is identified.
[0101] 3) Identify whether the user video data includes user hand features. If the user hand features are included, determine corresponding expression information based on the user hand features.
[0102] In this embodiment, after obtaining user video data, feature extraction is performed on the user video data, and it is identified whether the extracted features include user hand features. If user hand features are included, it indicates that the user may input through gestures, and the corresponding expression information can be determined based on the user hand features.
[0103] In specific implementation, an image recognition model that can recognize gestures (i.e., a gesture image recognition model) can be pre-trained. In actual application, the gesture image recognition model can be used to identify whether the extracted features include user hand features. If user hand features are included, the expression information included in the user video data can be identified.
[0104] 4) Identify whether the user hand video data includes user hand features. If the user hand features are included, determine whether the user hand features change continuously. If the user hand features change continuously, determine the corresponding semantics based on the user hand features, and determine the corresponding expression information based on the semantics.
[0105] In this embodiment, after obtaining user video data, feature extraction is performed on the user video data to identify whether the extracted features include user hand features. If the extracted features include user hand features, it is then determined whether the hand features in the user hand features change continuously. If the hand features change continuously, it indicates that the user may be inputting through sign language. The corresponding semantics are determined based on the hand features, and the corresponding expression information is determined based on the semantics.
[0106] In specific implementation, an image recognition model that can recognize sign language (i.e., a sign language image recognition model) can be pre-trained. In actual application, the sign language image recognition model can be used to identify whether the extracted features include user hand features. If the user hand features are included, it is then determined whether the hand features in the user hand features change continuously. If the hand features change continuously, the expression information included in the user video data is identified.
[0107] It is understood that in actual applications, corresponding image recognition models can be pre-trained for any input method (facial expression, lip reading, gesture, sign language), for example, facial expression image recognition models, lip reading image recognition models, gesture image recognition models, sign language image recognition models, etc. When user video data is collected, the user video data can be input into each image recognition model to obtain expression information recognized by each image recognition model, and the expression information recognized by each image recognition model is displayed as a candidate.
[0108] It is understood that when user video data is input into various image recognition models, if the user video data does not include features that a particular image recognition model can recognize, that image recognition model may not output the corresponding expression information for the user video data. For example, by simultaneously inputting user video data into a facial expression image recognition model, a lip reading image recognition model, a gesture image recognition model, and a sign language image recognition model, the expression information output by the facial expression image recognition model and the gesture image recognition model can be obtained.
[0109] Another method is to first classify the user video data when it is collected, and then identify the expression information included in the user video data based on the video classification result. Specifically, identifying the expression information included in the user video data includes:
[0110] 1) Classify user video data and obtain video classification results.
[0111] In this embodiment, when obtaining user video data, the user video data can be classified first to determine the specific type of the collected user video data, such as lip reading video data, facial expression video data, sign language video data, or gesture video data.
[0112] In a specific implementation, user video data can be classified to obtain a video classification result by: extracting user facial features and / or user hand features from the user video data; and determining a video classification result based on the user facial features and / or user hand features. That is, when the acquired user video data only includes user facial video data, the user facial features are extracted from the user facial video data, and the type of the user facial data video is determined based on the user facial features. When the acquired user video data only includes user hand video data, the user hand features are extracted from the user hand video data, and the type of the user hand data video is determined based on the user hand features. When the acquired user video data includes both user facial video data and user hand video data, the user facial features and user hand features are simultaneously extracted to determine the type of the user video data based on the user facial features and user hand features.
[0113] Specifically, determining the video classification result based on the user's facial features and / or the user's hand features includes:
[0114] When user video data includes user facial features, if the lip shape features within the user's facial features continuously change, the video classification result includes lip reading video data; otherwise, the video classification result includes facial expression video data. Specifically, if the lip shape features within the extracted user facial features continuously change, indicating a change in the user's mouth shape and a high probability that the user is using lip reading for input, the video classification result corresponding to the user video data at least includes lip reading video data. Continuous changes in lip shape features can mean that the acquired lip shape features continuously change over a predetermined period of time. If the lip shape features do not continuously change, indicating that the user may be using facial expressions for input, the video classification result corresponding to the user video data includes facial expression video data.
[0115] When user video data includes user hand features, if the hand features continuously change, the video classification result includes sign language video data; otherwise, the video classification result includes gesture video data. Specifically, if the extracted hand features continuously change, indicating that the user may be inputting via sign language, the video classification result for the user video data will at least include sign language video data. If the extracted hand features remain unchanged, indicating that the user may be inputting via gestures, the video classification result for the user video data will at least include gesture video data.
[0116] 2) Identify the expression information included in the user video data based on the video classification results.
[0117] In this embodiment, after obtaining the video classification result, the expression information included in the user video data can be identified based on the video classification result.
[0118] Specifically, when the video classification result includes facial expression video data, the corresponding expression information is determined based on the user's facial features in the facial expression video data; when the video classification result includes lip reading video data, the pronunciation is determined based on the lip shape features in the lip reading video data, and the corresponding expression information is determined based on the pronunciation; when the video classification result includes gesture video data, the corresponding expression information is determined based on the user's hand features included in the gesture video data; when the video classification result includes sign language video data, the corresponding semantics are determined based on the user's hand features in the sign language video data, and the corresponding expression information is determined based on the semantics. It should be noted that when a user video data corresponds to multiple video classification results, the expression information corresponding to each video classification result is obtained, and each expression information is then used as a candidate for the user to select.
[0119] In a specific implementation, image recognition models corresponding to each video type can be pre-trained based on the video type, for example, facial expression image recognition models, lip reading image recognition models, gesture image recognition models, and sign language image recognition models. After determining the video classification results corresponding to the user's video data, the corresponding image recognition model is determined based on the video classification results. This image recognition model is then used to identify the expression information included in the user's video data, thereby improving recognition efficiency and accuracy, providing users with precise expression information and enhancing the user experience.
[0120] That is, when the video classification result of the user video data includes lip reading video data, the lip reading image recognition model is used to identify the expression information included in the user video data; when the video classification result of the user video data includes facial expression video data, the facial expression image recognition model is used to identify the expression information included in the corresponding user video data; when the video classification result of the user video data includes sign language video data, the sign language image recognition model is used to identify the expression information included in the user video data; when the video classification result of the user video data includes gesture video data, the gesture image recognition model is used to identify the expression information included in the user video data.
[0121] It is understandable that when collecting user video data, the user video data may include the above-mentioned multiple types of video data at the same time. Therefore, when performing video classification, there may be corresponding multiple video classification results. When multiple video classification results are obtained, the image recognition model corresponding to each video classification result is used for recognition, so that multiple expression information can be obtained, and then the multiple expression information is displayed to the user for the user to select the expression information to be expressed. For example Figure 3b As shown, the user video data includes both facial and hand video data. By extracting the user's facial and hand features, the user's expressive information can be determined based on the facial and hand features. For example, if the expressive information obtained from the user's facial features is "I'm so tired, I don't want to move" or "So awkward," and the expressive information obtained from the user's hand features is "It's none of my business" or "What's the deal?", then all of the identified expressive information will be displayed as candidate items, allowing the user to select the text they want to express.
[0122] Based on the above method embodiment, the present application also provides an input device, which will be described below with reference to the accompanying drawings.
[0123] See also Figure 4 , which is a structural diagram of an input device provided in an embodiment of the present application, such as Figure 4 As shown, the device includes:
[0124] The acquisition unit 401 is configured to acquire user video data, wherein the user video data includes user facial video data and / or user hand video data;
[0125] an identification unit 402, configured to identify expression information included in the user video data;
[0126] The display unit 403 is configured to display the expression information as a candidate item.
[0127] In a possible implementation, the identification unit includes:
[0128] A classification subunit, configured to classify the user video data and obtain a video classification result;
[0129] The identification subunit is used to identify the expression information included in the user video data according to the video classification result.
[0130] In a possible implementation, the recognition subunit is specifically configured to, when the video classification result includes facial expression video data, determine corresponding expression information based on the user's facial features in the facial expression video data;
[0131] When the video classification result includes lip reading video data, determining pronunciation according to lip shape features in the lip reading video data, and determining corresponding expression information according to the pronunciation;
[0132] When the video classification result includes gesture video data, determining corresponding expression information according to user hand features included in the gesture video data;
[0133] When the video classification result includes sign language video data, corresponding semantics are determined according to user hand features in the sign language video data, and corresponding expression information is determined according to the semantics.
[0134] In a possible implementation, the classification subunit includes:
[0135] an extraction subunit, configured to extract user facial features and / or user hand features from the user video data;
[0136] A determination subunit is used to determine the video classification result based on the user's facial features and / or the user's hand features.
[0137] In a possible implementation, the determining subunit is specifically configured to, when the user video data includes the user facial features, if the lip shape features in the user facial features continuously change, the video classification result includes lip reading video data; otherwise, the video classification result includes facial expression video data;
[0138] When the user video data includes the user hand feature, if the user hand feature changes continuously, the video classification result includes sign language video data; otherwise, the video classification result includes gesture video data.
[0139] In a possible implementation, the identification unit includes:
[0140] a first recognition subunit, configured to identify whether the user video data includes user facial features, and if so, determine corresponding expression information based on the user facial features;
[0141] a second recognition subunit, configured to identify whether the user video data includes user facial features, and if so, determine whether a lip shape feature in the user facial features continuously changes; if so, determine a pronunciation based on the lip shape feature, and determine corresponding expression information based on the pronunciation;
[0142] a third identification subunit, configured to identify whether the user video data includes a user hand feature, and if so, determine corresponding expression information according to the user hand feature;
[0143] The fourth identification subunit is used to identify whether the user hand video data includes user hand features. If the user hand features are included, determine whether the user hand features change continuously. If the user hand features change continuously, determine the corresponding semantics based on the user hand features, and determine the corresponding expression information based on the semantics.
[0144] In a possible implementation, the expression information includes one or more of expression text, image, and expression.
[0145] In a possible implementation, the expression text includes a standard expression text and / or a custom expression text;
[0146] When the expression text includes a custom expression text, the display unit is specifically configured to preferentially display the custom expression text among the candidate items.
[0147] It should be noted that the implementation of each unit in this embodiment can refer to the above method embodiment, and this embodiment will not be repeated here.
[0148] Figure 5 FIG2 shows a block diagram of an input device 600. For example, the device 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, or the like.
[0149] Reference Figure 5 , the device 600 may include one or more of the following components: a processing component 602 , a memory 604 , a power component 606 , a multimedia component 608 , an audio component 610 , an input / output (I / O) interface 69 , a sensor component 614 , and a communication component 616 .
[0150] The processing component 602 generally controls the overall operation of the device 600, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 602 may include one or more modules to facilitate interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate interaction between the multimedia component 606 and the processing component 602.
[0151] The memory 604 is configured to store various types of data to support operations on the device 600. Examples of such data include instructions for any application or method operating on the device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0152] The power supply component 606 provides power to the various components of the device 600. The power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 600.
[0153] The multimedia component 608 includes a screen that provides an output interface between the device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0154] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC), which is configured to receive external audio signals when the device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 also includes a speaker for outputting audio signals.
[0155] The I / O interface provides an interface between the processing component 602 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0156] The sensor assembly 614 includes one or more sensors for providing various aspects of the status assessment of the device 600. For example, the sensor assembly 614 can detect the open / closed state of the device 600, the relative positioning of components, such as the display and keypad of the device 600. The sensor assembly 614 can also detect changes in the position of the device 600 or a component of the device 600, the presence or absence of user contact with the device 600, the orientation or acceleration / deceleration of the device 600, and temperature changes of the device 600. The sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0157] The communication component 616 is configured to facilitate wired or wireless communication between the device 600 and other devices. The device 600 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0158] In an exemplary embodiment, the apparatus 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the following method:
[0159] Collecting user video data, where the user video data includes user facial video data and / or user hand video data;
[0160] Identifying expression information included in the user video data;
[0161] The expression information is displayed as a candidate.
[0162] Optionally, the identifying expression information included in the user video data includes:
[0163] Classifying the user video data to obtain a video classification result;
[0164] The expression information included in the user video data is identified according to the video classification result.
[0165] Optionally, the identifying, according to the video classification result, expression information included in the user video data includes:
[0166] When the video classification result includes facial expression video data, determining corresponding expression information based on the user's facial features in the facial expression video data;
[0167] When the video classification result includes lip reading video data, determining pronunciation according to lip shape features in the lip reading video data, and determining corresponding expression information according to the pronunciation;
[0168] When the video classification result includes gesture video data, determining corresponding expression information according to user hand features included in the gesture video data;
[0169] When the video classification result includes sign language video data, corresponding semantics are determined according to user hand features in the sign language video data, and corresponding expression information is determined according to the semantics.
[0170] Optionally, classifying the user video data to obtain a video classification result includes:
[0171] Extracting user facial features and / or user hand features from the user video data;
[0172] The video classification result is determined based on the user's facial features and / or the user's hand features.
[0173] Optionally, determining the video classification result according to the user's facial features and / or the user's hand features includes:
[0174] When the user video data includes the user's facial features, if the lip shape features in the user's facial features continuously change, the video classification result includes lip reading video data; otherwise, the video classification result includes facial expression video data;
[0175] When the user video data includes the user hand feature, if the user hand feature changes continuously, the video classification result includes sign language video data; otherwise, the video classification result includes gesture video data.
[0176] Optionally, the identifying expression information included in the user video data includes:
[0177] Identifying whether the user video data includes user facial features, and if so, determining corresponding expression information based on the user facial features;
[0178] Identifying whether the user video data includes user facial features, and if so, determining whether a lip shape feature in the user facial features continuously changes, and if so, determining a pronunciation based on the lip shape feature, and determining corresponding expression information based on the pronunciation;
[0179] Identifying whether the user video data includes a user hand feature, and if so, determining corresponding expression information based on the user hand feature;
[0180] Identify whether the user hand video data includes user hand features; if so, determine whether the user hand features continuously change; if so, determine corresponding semantics based on the user hand features, and determine corresponding expression information based on the semantics.
[0181] Optionally, the expression information includes one or more of text, images, and expressions.
[0182] Optionally, the expression text includes a standard expression text and / or a custom expression text;
[0183] When the expression text includes a custom expression text, displaying the expression information as a candidate item includes:
[0184] The custom expression text is displayed preferentially among the candidate items.
[0185] A non-transitory computer-readable storage medium, wherein instructions in the storage medium are used by a processor of a mobile terminal to collect user video data, wherein the user video data includes user facial video data and / or user hand video data;
[0186] Identifying expression information included in the user video data;
[0187] The expression information is displayed as a candidate.
[0188] Optionally, the identifying textual expression information included in the user video data includes:
[0189] Classifying the user video data to obtain a video classification result;
[0190] The expression text information included in the user video data is identified according to the video classification result.
[0191] Optionally, the identifying, according to the video classification result, expression information included in the user video data includes:
[0192] When the video classification result includes facial expression video data, determining corresponding expression information based on the user's facial features in the facial expression video data;
[0193] When the video classification result includes lip reading video data, determining pronunciation according to lip shape features in the lip reading video data, and determining corresponding expression information according to the pronunciation;
[0194] When the video classification result includes gesture video data, determining corresponding expression information according to user hand features included in the gesture video data;
[0195] When the video classification result includes sign language video data, corresponding semantics are determined according to user hand features in the sign language video data, and corresponding expression information is determined according to the semantics.
[0196] Optionally, classifying the user video data to obtain a video classification result includes:
[0197] Extracting user facial features and / or user hand features from the user video data;
[0198] The video classification result is determined based on the user's facial features and / or the user's hand features.
[0199] Optionally, determining the video classification result according to the user's facial features and / or the user's hand features includes:
[0200] When the user video data includes the user's facial features, if the lip shape features in the user's facial features continuously change, the video classification result includes lip reading video data; otherwise, the video classification result includes facial expression video data;
[0201] When the user video data includes the user hand feature, if the user hand feature changes continuously, the video classification result includes sign language video data; otherwise, the video classification result includes gesture video data.
[0202] Optionally, the identifying expression information included in the user video data includes:
[0203] Identifying whether the user video data includes user facial features, and if so, determining corresponding expression information based on the user facial features;
[0204] Identifying whether the user video data includes user facial features, and if so, determining whether a lip shape feature in the user facial features continuously changes, and if so, determining a pronunciation based on the lip shape feature, and determining corresponding expression information based on the pronunciation;
[0205] Identifying whether the user video data includes a user hand feature, and if so, determining corresponding expression information based on the user hand feature;
[0206] Identify whether the user hand video data includes user hand features; if so, determine whether the user hand features continuously change; if so, determine corresponding semantics based on the user hand features, and determine corresponding expression information based on the semantics.
[0207] Optionally, the expression information includes one or more of text, images, and expressions.
[0208] Optionally, the expression text includes a standard expression text and / or a custom expression text;
[0209] When the expression text includes a custom expression text, displaying the expression information as a candidate item includes:
[0210] The custom expression text is displayed preferentially among the candidate items.
[0211] Figure 6 7 is a schematic diagram of the structure of the server in an embodiment of the present invention. The server 700 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 722 (for example, one or more processors) and memory 732, and one or more storage media 730 (for example, one or more mass storage devices) for storing application programs 742 or data 744. Among them, the memory 732 and the storage medium 730 can be temporary storage or permanent storage. The program stored in the storage medium 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 722 can be configured to communicate with the storage medium 730 to execute a series of instruction operations in the storage medium 730 on the server 700.
[0212] The terminal 700 may also include one or more power supplies 726, one or more wired or wireless network interfaces 750, one or more input and output interfaces 756, one or more keyboards 756, and / or one or more operating systems 741, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0213] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0214] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0215] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0216] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0217] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An input method, characterized in that: The method comprises: When the user is unable to use voice input and pinyin input, the input method client collects user video data, wherein the user video data includes user facial video data and user hand video data, wherein the user facial video data includes lip language video data and / or facial expression video data, and the user hand video data includes sign language video data and / or gesture video data; Classifying the user video data to obtain a video classification result; Identify, according to an image recognition model corresponding to the video classification result, expression information included in the user video data, the expression information including multiple of expression text, images, and expressions, the expression text including standard text and / or custom expression text, the expression being a dynamic expression image or a static expression image, and the image recognition model including a facial expression image recognition model, a lip reading image recognition model, a gesture image recognition model, and a sign language image recognition model; displaying all of the expression information as candidates; Wherein, when the expression text includes a custom expression text, displaying the expression information as a candidate item includes: Prioritizing display of the custom expression text among the candidate items; The identifying expression information included in the user video data includes: Using a facial expression image recognition model to identify whether the user video data includes user facial features, and if the user facial features are included, determining corresponding expression information based on the user facial features; Using the lip reading image recognition model to identify whether the user video data includes user facial features, if the user facial features are included, determining whether lip shape features in the user facial features continuously change, if the lip shape features continuously change, determining pronunciation based on the lip shape features, and determining corresponding expression information based on the pronunciation; using the gesture image recognition model to identify whether the user video data includes user hand features, and if so, determining corresponding expression information based on the user hand features; The sign language image recognition model is used to identify whether the user hand video data includes user hand features. If the user hand features are included, it is determined whether the user hand features change continuously. If the user hand features change continuously, the corresponding semantics are determined according to the user hand features, and the corresponding expression information is determined according to the semantics.
2. The method according to claim 1, characterized in that The classifying the user video data to obtain a video classification result includes: Extracting user facial features and / or user hand features from the user video data; The video classification result is determined based on the user's facial features and / or the user's hand features.
3. The method according to claim 2, characterized in that Determining the video classification result according to the user's facial features and / or the user's hand features includes: When the user video data includes the user's facial features, if the lip shape features in the user's facial features continuously change, the video classification result includes lip reading video data; otherwise, the video classification result includes facial expression video data; When the user video data includes the user hand feature, if the user hand feature changes continuously, the video classification result includes sign language video data; otherwise, the video classification result includes gesture video data.
4. An input device, characterized in that: The device comprises: a collection unit, configured to collect user video data by the input method client when the user is unable to use voice input and pinyin input, wherein the user video data includes user facial video data and user hand video data, wherein the user facial video data includes lip language video data and / or facial expression video data, and the user hand video data includes sign language video data and / or gesture video data; an identification unit, configured to identify expression information included in the user video data; the expression information includes multiple expression texts, images, and expressions, and the expression texts include standard expression texts and / or custom expression texts; A display unit, configured to display all the expression information as candidate items; The identification unit includes: A classification subunit, configured to classify the user video data and obtain a video classification result; an identification subunit, configured to identify expression information included in the user video data according to an image recognition model corresponding to the video classification result, wherein the expression is a dynamic expression image or a static expression image, and the image recognition model includes a facial expression image recognition model, a lip reading image recognition model, a gesture image recognition model, and a sign language image recognition model; The display unit is specifically configured to, when the expression text includes a custom expression text, preferentially display the custom expression text among the candidate items; The identification unit includes: A first recognition subunit is configured to use a facial expression image recognition model to identify whether the user video data includes user facial features, and if so, determine corresponding expression information based on the user facial features; a second recognition subunit, configured to use the lip reading image recognition model to identify whether the user video data includes user facial features; if so, determine whether lip shape features in the user facial features continuously change; if so, determine pronunciation based on the lip shape features, and determine corresponding expression information based on the pronunciation; a third recognition subunit, configured to use the gesture image recognition model to identify whether the user video data includes user hand features, and if so, determine corresponding expression information based on the user hand features; The fourth identification subunit is used to use the sign language image recognition model to identify whether the user hand video data includes user hand features; if the user hand features are included, determine whether the user hand features change continuously; if the user hand features change continuously, determine the corresponding semantics based on the user hand features, and determine the corresponding expression information based on the semantics.
5. The device according to claim 4, characterized in that The classification subunit includes: an extraction subunit, configured to extract user facial features and / or user hand features from the user video data; A determination subunit is used to determine the video classification result based on the user's facial features and / or the user's hand features.
6. The device according to claim 5, characterized in that The determining subunit is specifically configured to, when the user video data includes the user facial features, determine that if the lip shape features in the user facial features continuously change, the video classification result includes lip reading video data; otherwise, the video classification result includes facial expression video data; When the user video data includes the user hand feature, if the user hand feature changes continuously, the video classification result includes sign language video data; Otherwise, the video classification result includes gesture video data.
7. A device for input, characterized in that The system includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: When the user is unable to use voice input and pinyin input, the input method client collects user video data, wherein the user video data includes user facial video data and user hand video data, wherein the user facial video data includes lip language video data and / or facial expression video data, and the user hand video data includes sign language video data and / or gesture video data; Classifying the user video data to obtain a video classification result; Identify, according to an image recognition model corresponding to the video classification result, expression information included in the user video data, the expression information including multiple of expression text, images, and expressions, the expression text including standard text and / or custom expression text, the expression being a dynamic expression image or a static expression image, and the image recognition model including a facial expression image recognition model, a lip reading image recognition model, a gesture image recognition model, and a sign language image recognition model; displaying all of the expression information as candidates; Wherein, when the expression text includes a custom expression text, displaying the expression information as a candidate item includes: Prioritizing display of the custom expression text among the candidate items; The identifying expression information included in the user video data includes: Using a facial expression image recognition model, identifying whether the user video data includes user facial features, and if the user facial features are included, determining corresponding expression information based on the user facial features; Using the lip reading image recognition model to identify whether the user video data includes user facial features, if the user facial features are included, determining whether lip shape features in the user facial features continuously change, if the lip shape features continuously change, determining pronunciation based on the lip shape features, and determining corresponding expression information based on the pronunciation; using the gesture image recognition model to identify whether the user video data includes user hand features, and if so, determining corresponding expression information based on the user hand features; The sign language image recognition model is used to identify whether the user hand video data includes user hand features. If the user hand features are included, it is determined whether the user hand features change continuously. If the user hand features change continuously, the corresponding semantics are determined according to the user hand features, and the corresponding expression information is determined according to the semantics.
8. A computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause a device to perform the input method according to any one of claims 1 to 5.