A method and apparatus for selecting an interactive action
Patent Information
- Application Number
- CN202210043059.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-01-14
AI Technical Summary
而在实践中发现,基于用户习惯选择的交互动作有时候会出现设备对其识别的准确率较低的问题,从而给用户带来不好的体验
[0026]本公开实施例提供的交互动作的选择方法及装置,通过将所述交互动作集合中的各交互动作分别对应的使用度参数和识别度参数进行融合,得到所述交互动作对应的融合参数,从而依据所述融合参数确定所述目标交互动作的方法,使得所述目标交互动作既符合大多数用户的使用习惯,给用户带来较好的体验,又由于在交互动作的选择过程中考虑了设备对于交互动作的识别准确率的因素,使得最终确定的交互动作的识别准确率也较高。
Smart Images

Figure CN116483193B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of human-computer interaction technology, and in particular to a method and apparatus for selecting interactive actions. Background Technology
[0002] With the continuous improvement of computer hardware and software capabilities and the increasing demand for human-computer interaction, the selection of interactive actions has received more and more attention. Interactive actions are typically performed by making predefined movements in front of a camera, which are then recognized by machine learning models. This allows users to control devices using motion interaction when it is inconvenient to directly touch the control panel. This is commonly used in AR (Augmented Reality), VR (Virtual Reality), smart TVs, and in-vehicle terminals in cars.
[0003] In related technologies, the choice of which interactive gestures to use for device control involves the problem of gesture selection. For example, taking television control via interactive gestures, there are many available gestures, including the V-sign, OK sign, and swiping forward. However, television control typically doesn't require such a large number of gestures. The problem lies in selecting five gestures from this pool for television control. In one example of related technologies, selection is often based on user habits; for instance, if most users prefer the V-sign, it's used as the control gesture. However, in practice, it has been found that interaction gestures selected based on user habits sometimes result in lower device recognition accuracy, leading to a poor user experience. Summary of the Invention
[0004] In view of this, the present disclosure provides a method and apparatus for selecting interactive actions.
[0005] Specifically, the embodiments of this disclosure are implemented through the following technical solutions:
[0006] According to a first aspect of this application, a method for selecting an interactive action is proposed, the method comprising:
[0007] Obtain the usage parameter corresponding to each interactive action in the set of interactive actions, and the usage parameter is used to characterize the user's usage preference for the interactive action;
[0008] Obtain the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions, and the recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized;
[0009] For each of the interactive actions, the usage parameter and the recognition parameter are fused to obtain the fused parameter corresponding to the interactive action;
[0010] In response to the fusion parameters meeting preset parameter selection conditions, the interactive action corresponding to the fusion parameters is determined as the target interactive action.
[0011] According to a second aspect of this disclosure, an interactive action selection device is provided, the device comprising:
[0012] The first acquisition module is used to acquire the usage parameters corresponding to each interactive action in the set of interactive actions, and the usage parameters are used to characterize the user's usage preference for the interactive action.
[0013] The second acquisition module is used to acquire the recognition degree parameter corresponding to each interactive action in the set of interactive actions, and the recognition degree parameter is used to characterize the accuracy of the interactive action being recognized.
[0014] The fusion module is used to fuse the usage parameter and the recognition parameter for each interaction action to obtain the fusion parameter corresponding to the interaction action.
[0015] The target interaction action determination module is used to determine the interaction action corresponding to the fusion parameters as the target interaction action in response to the fusion parameters meeting the preset parameter selection conditions.
[0016] According to a third aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the method described in any embodiment of the first aspect above.
[0017] According to a fourth aspect of this disclosure, an electronic device is provided, comprising:
[0018] processor;
[0019] Memory used to store processor-executable instructions;
[0020] The processor is configured as follows:
[0021] Obtain the usage parameter corresponding to each interactive action in the set of interactive actions, and the usage parameter is used to characterize the user's usage preference for the interactive action;
[0022] Obtain the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions, and the recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized;
[0023] For each of the interactive actions, the usage parameter and the recognition parameter are fused to obtain the fused parameter corresponding to the interactive action;
[0024] In response to the fusion parameters meeting preset parameter selection conditions, the interactive action corresponding to the fusion parameters is determined as the target interactive action.
[0025] The technical solutions provided in this disclosure may have the following beneficial effects:
[0026] The interactive action selection method and apparatus provided in this disclosure fuse the usage parameters and recognition parameters corresponding to each interactive action in the interactive action set to obtain the fusion parameters corresponding to the interactive action, and then determine the target interactive action based on the fusion parameters. This method ensures that the target interactive action not only conforms to the usage habits of most users and provides a better user experience, but also has a high recognition accuracy because the device's recognition accuracy is considered during the selection process.
[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0028] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0029] Figure 1 This is a flowchart illustrating an interactive action selection method according to an exemplary embodiment of the present disclosure;
[0030] Figure 2 This is a flowchart illustrating another method for selecting an interactive action according to an exemplary embodiment of the present disclosure;
[0031] Figure 3 This is a schematic diagram of the structure of an interactive action selection device according to an exemplary embodiment of the present disclosure;
[0032] Figure 4 This is a schematic diagram of the structure of another interactive action selection device according to an exemplary embodiment of the present disclosure;
[0033] Figure 5 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0035] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0036] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0037] Existing methods for selecting interactive actions are typically based on user preferences, the device's ability to recognize the action, or device limitations. For example, in a scenario where a user wants to turn on a TV channel and has already selected it, they might prefer to use a "push forward" gesture to confirm. However, when a user chooses to use this gesture to confirm, the limitations of the TV and the accuracy of the TV's recognition of this action are not considered. This is because recognizing the "push forward" gesture requires a depth camera on the TV, and even with a depth camera, the accuracy will be low if the user is far from it. In other words, existing technologies consider user preferences, the device's accuracy in recognizing the action, and the device's limitations in a relatively isolated way when selecting interactive actions, and lack scientific basis.
[0038] In view of this, the present disclosure provides a method and apparatus for selecting interactive actions.
[0039] The method for selecting interactive actions is used to select a subset of interactive actions from a set of interactive actions to be selected as target interactive actions.
[0040] The set of interactive actions to be selected is a collection of interactive actions that can be used to control the device in a specific scenario. For example, in a scenario where a television is controlled via interactive actions, the interactive actions that can be used to control the television include a large number of actions such as the V gesture, the OK gesture, and the push-forward gesture. All of these interactive actions can constitute the set of interactive actions to be selected.
[0041] The target interactive action is an interactive action selected from the set of interactive actions for controlling the device in the specific scenario. For example, in a scenario where a television is controlled via an interactive action, the forward push gesture can be selected from a large number of gestures such as the V gesture, OK gesture, and forward push gesture as the interactive action to control the television. The selected forward push gesture is the target interactive action.
[0042] The method for selecting interactive actions according to embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0043] Figure 1 This is a flowchart illustrating a method for selecting an interactive action according to an exemplary embodiment of this disclosure. Figure 1 As shown, the exemplary embodiment method may include the following processes:
[0044] Step S101: Obtain the usage parameter corresponding to each interactive action in the set of interactive actions. The usage parameter is used to characterize the user's usage preference for the interactive action.
[0045] The interactive actions can include static and dynamic interactive actions, and both static and dynamic interactive actions can include gesture interactive actions, human posture interactive actions, etc. This disclosure does not limit the specific form of the interactive actions.
[0046] In one optional example, for each interaction action in the set of interaction actions, selection information corresponding to the interaction action can be received, including the number of people who selected the interaction action. The number of selections is then tallied to obtain a usage parameter corresponding to the interaction action.
[0047] For example, gestures such as the V-sign, OK-sign, and push-forward gesture, which are part of the set of interactive actions for controlling a TV, can be displayed on the screen. 100 users can then choose their preferred interactive action for controlling the TV (for example, each user can choose one action), thus obtaining the number of users who selected each action. For instance, after the 100 users have made their selections, the number of users who selected each action can be counted, showing that 18 users chose the V-sign and 20 users chose the OK-sign, etc. Furthermore, the proportion of users who selected each action relative to the total number of users can be used as a usage parameter for that action.
[0048] For example, the usage parameter of the V gesture is 0.18, which is the number of people who selected the V gesture (18) divided by the total number of people (100); the usage parameter of the OK gesture is 0.2, which is the number of people who selected the OK gesture (20) divided by the total number of people (100). Furthermore, the sum of the usage parameters of all interactive actions displayed on the screen for controlling the TV is 1.
[0049] In one optional example, for each interaction action in the set of interaction actions, the frequency with which the interaction action is selected over a period of time can be received. The frequency of use is then statistically analyzed to obtain a usage parameter corresponding to the interaction action.
[0050] For example, the set of interactive actions for controlling the TV, including the V-sign, OK-sign, and push-forward gesture, can be displayed on the screen, allowing users to select their favorite actions. The selection results are statistically analyzed over a period of time to obtain the number of people who selected each interactive action during that period. Furthermore, the frequency of selection for each interactive action can be obtained, i.e., the usage parameter of each interactive action. For example, if 200 selections were made on the screen over two days, and the V-sign was selected 24 times, then the frequency of V-sign selection, i.e., the usage parameter of V-sign, is 0.12; if the OK-sign was selected 36 times, then the frequency of OK-sign selection, i.e., the usage parameter of OK-sign, is 0.18, and so on. Similarly, under this statistical method, the sum of the usage parameters of all interactive actions displayed on the screen for controlling the TV is also 1. In an optional example, the selection information can include both the number of people who selected the interactive action and the frequency of the interactive action. In this case, the number of people who selected the interactive action and the frequency of the interactive action can be combined according to the following formula:
[0051] Z = aE + bY (1)
[0052] Wherein, Z represents the usage parameter, a represents the weight corresponding to the number of people E who selected the interactive action, and b represents the weight corresponding to the frequency Y of the interactive action.
[0053] In one alternative example, when a user selects an interactive action displayed on the screen, they can choose one of their favorite interactive actions when controlling the TV, or they can choose two interactive actions that they prefer to use. This disclosure does not impose any specific limitations on this.
[0054] Furthermore, in one example, the method of this embodiment can be executed by an interactive action selection device, in which case the aforementioned usage parameters can be obtained by the interactive action selection device itself.
[0055] For example, as described above, the interactive action selection device can display various actions, such as the V gesture, from the set of interactive actions used to control the television on the screen and receive the user's selection information.
[0056] For example, the interactive action selection device can also push multiple actions included in the interactive action set to each user's own mobile phone, so that the user can select the action on their own mobile phone. The interactive action selection device receives the selection information returned by the user's mobile phone and then performs statistics on the selection information.
[0057] For example, the aforementioned usage parameters can also be obtained by statistics from other devices besides the interaction action selection device, and then transmitted to the interaction action selection device.
[0058] Step S102: Obtain the recognition parameter corresponding to each interactive action in the set of interactive actions. The recognition parameter is used to characterize the accuracy of the interactive action being recognized.
[0059] In one optional example, the recognition parameter corresponding to each interactive action is obtained by fusing the recognition rate and the false judgment rate corresponding to each interactive action.
[0060] In an optional example, a category recognizer can be used to obtain the recognition rate and false positive rate of the interaction action.
[0061] For example, 100 different forward pushing gestures can be input into the category recognizer, and the number of forward pushing gestures recognized by the category recognizer can be counted. For example, if the number of forward pushing gestures recognized by the category recognizer is 90, then the recognition rate of the forward pushing gesture is 90%.
[0062] For example, 100 non-forward push gestures can be input into the category recognizer, and then the number of forward push gestures recognized by the category recognizer can be counted. For example, if the category recognizer recognizes 10 forward push gestures, then the misjudgment rate of the forward push gesture is 10%.
[0063] Specifically, for each interaction action, after obtaining the recognition rate and false positive rate corresponding to the interaction action, the recognition rate and false positive rate of the interaction action can be fused to obtain the recognition degree parameter corresponding to the interaction action.
[0064] In an optional example, the false positive rate and recognition rate of the interactive action can be fused according to the following formula to obtain the recognition parameter corresponding to the interactive action:
[0065] X R =W r ×X recognize +W wj ×X wrongJudge (2)
[0066] Wherein, X R The W represents the recognition parameter of the interactive action. r The recognition rate X represents the interaction action. recognize The corresponding weight, the W wj The false positive rate X represents the interaction action. wrongJudge The corresponding weights. The W r and the W wj These can be called empirical values, and their numerical values satisfy the following formula:
[0067] W r +W wj =1 (3)
[0068] Furthermore, when fusing the recognition rate and the false positive rate, a higher recognition rate indicates a higher accuracy represented by the recognition degree parameter, and a lower false positive rate indicates a higher accuracy represented by the recognition degree parameter. Therefore, the W r The value should be greater than 0, while the W mentioned above... wj The value should be less than 0. W r and W wj The specific value depends on factors such as the maturity of the technical solution, the operating environment of the algorithm, and the product settings, which will not be discussed in detail here.
[0069] Furthermore, the above formula (1) is only an exemplary way of fusing the false positive rate and the recognition rate. In specific implementation, it is not limited to this and other methods can also be used for fusion.
[0070] In an optional example, the recognition parameter corresponding to each interactive action can be determined solely by the recognition rate or the false positive rate corresponding to each interactive action.
[0071] Step S103: For each interaction action, the usage parameter and the recognition parameter are fused to obtain the fusion parameter corresponding to the interaction action.
[0072] In an optional example, the usage parameter and recognition parameter of the interactive action can be fused according to the following formula to obtain the fused parameter corresponding to the interactive action:
[0073] X = W r ×X R +W u ×X user (4)
[0074] Wherein, X represents the fusion parameter corresponding to the interactive action, and W... r The recognition rate parameter X represents the interaction action. R The corresponding weight, the W u The usage parameter X represents the degree of use of the interactive action. user The corresponding weights. The W r and the W u These can be called empirical values, and their numerical values also satisfy formula (3).
[0075] In an optional example, the W r and the W u The value of can be 0.5, indicating that the usage parameter and the recognition parameter have the same impact on the fusion parameter.
[0076] Similarly, the above formula (4) is only an exemplary way of fusing the usage parameter and recognition parameter of the interactive action. In specific implementation, it is not limited to this and other methods can also be used for fusion.
[0077] The interactive action selection method in this example fuses the usage and recognition parameters corresponding to each interactive action in the interactive action set to obtain a fusion parameter. The interactive actions in the interactive action set are then selected based on the fusion parameter, making the selection of interactive actions more scientifically meaningful.
[0078] Step S104: In response to the fusion parameters meeting the preset parameter selection conditions, the interactive action corresponding to the fusion parameters is determined as the target interactive action.
[0079] The parameter selection criteria are used to filter the fusion parameters. When the fusion parameter meets the parameter selection criteria, the interactive action corresponding to the fusion parameter is determined as the target interactive action.
[0080] In an optional example, the parameter selection condition may be to sort the fusion parameters corresponding to each interactive action in the set of interactive actions, and determine the interactive action corresponding to the fusion parameter ranked at a preset position as the target interactive action based on the sorted fusion parameters.
[0081] For example, the obtained fusion parameters can be 0.8, 0.4, 0.5, and 0.6 respectively. These fusion parameters are then sorted in descending order to obtain 0.8, 0.6, 0.5, and 0.4. Based on the sorted fusion parameters, the interaction actions corresponding to the first two fusion parameters are determined as the target interaction actions; that is, the interaction actions corresponding to 0.8 and 0.6 are taken as the target interaction actions.
[0082] In an optional example, the fusion parameter can carry a label indicating the correspondence between the fusion parameter and the interactive action. After the above sorting operation is performed and the fusion parameters are determined to be in the preset order, the target interactive action can be determined based on the determined fusion parameters and the labels carried by the fusion parameters.
[0083] In an optional example, the parameter selection condition may be to determine the interactive action corresponding to the fusion parameter as the target interactive action in response to the value of the fusion parameter exceeding a preset parameter threshold.
[0084] For example, the parameter threshold can be preset to 0.5. When the value of the fusion parameter exceeds 0.5, the interactive action corresponding to the fusion parameter is determined as the target interactive action.
[0085] In this example, a threshold can also be set for the number of target interactive actions. When the number of target interactive actions obtained by the above parameter threshold method exceeds the threshold, the method of sorting the fusion parameters and selecting the interactive actions corresponding to the fusion parameters with a preset number of digits can be combined to further filter and obtain a suitable number of target interactive actions.
[0086] The interactive action selection method provided in this disclosure fuses the usage parameters and recognition parameters corresponding to each interactive action in the interactive action set to obtain the fusion parameters corresponding to the interactive action. The method then determines the target interactive action based on the fusion parameters. This method ensures that the target interactive action not only conforms to the usage habits of most users and provides a better user experience, but also has a high recognition accuracy because the device's recognition accuracy is considered during the selection process.
[0087] Figure 2 This is a flowchart illustrating another method for selecting interactive actions according to an exemplary embodiment. In the description of this embodiment, steps identical to those in any of the foregoing embodiments will be briefly described and will not be detailed further; however, specific details can be found in any of the foregoing embodiments. In the foregoing embodiments, a target interactive action is selected from a set of interactive actions to be selected. This embodiment also describes the process of obtaining this set of interactive actions. The set can be composed of multiple interactive actions with certain differences from an initial large number of interactive actions, thereby reducing the misjudgment rate during interactive action recognition. For example... Figure 2 As shown, the method in this embodiment may include the following processes:
[0088] Step S201: Obtain an initial image set including multiple action images.
[0089] In one optional example, motion images usable in a specific scenario can be pre-obtained, wherein each motion image includes an interactive action. The motion images including the interactive action usable in the specific scenario constitute the initial image set.
[0090] For example, when determining interactive actions to control a car, motion images containing human postures should not be acquired, as these are unsuitable for use in the confined space of a car. Instead, motion images containing gestures (exemplary examples) should be acquired, such as OK gestures, V gestures, etc.
[0091] Step S202: Extract features from each action image to obtain the feature vector corresponding to the interaction action.
[0092] In an optional example, the motion image can be input into a category recognizer, wherein the category recognizer is used to identify the category of the interactive action in the motion image, and the category recognizer includes a feature extraction module and a category output module.
[0093] The feature extraction module may include convolutional layers, non-linear activation layers, and pooling layers. The input to the feature extraction module is the action image, and the output is the feature vector. That is, the feature extraction module is used to extract features from the action image to obtain feature vectors corresponding to the interactive actions in the action image.
[0094] The category output module may include a fully connected layer. The input to the category output module is the feature vector output by the feature extraction module, and the output is the category of the interactive action. That is, the category output module is used to obtain the category of the interactive action in the action image based on the feature vector output by the feature extraction module. Therefore, in this step, during the processing of the action image using the category recognizer, the image does not pass through the category output module of the category recognizer, but rather is output after passing through the feature extraction module of the category recognizer, corresponding to the interactive action contained in the action image.
[0095] In an optional example, the action image can be preprocessed before being input into the category recognizer. Specifically, the action image can be cropped to obtain the portion containing the interactive action. For example, before inputting an action image of someone making an OK gesture into the category recognizer, the portion representing the OK gesture can be cropped out of the image before inputting it into the category recognizer. This reduces irrelevant information in the action image, enhances the detectability of relevant information, and simplifies the image data input into the category recognizer to the greatest extent possible.
[0096] Step S203: Obtain the category spacing between each interactive action based on the feature vectors corresponding to the interactive actions in each action image.
[0097] The category spacing is used to measure the similarity between different interactive actions.
[0098] In an optional example, the Euclidean distance algorithm can be used to calculate the class spacing between the feature vectors corresponding to the interactive actions in each action image.
[0099] Specifically, for two n-dimensional vectors a(X) 11 X 12 ..., X 1n ) and b(X 21 X 22 ..., X 2n The Euclidean distance between vector a and vector b is:
[0100]
[0101] Common methods for calculating the class distance between feature vectors corresponding to each interactive action in the images include Manhattan distance, Chebyshev distance, Minkowski distance, and standardized Euclidean distance. This disclosure does not specifically limit the method used to calculate the class distance between feature vectors corresponding to each interactive action.
[0102] Step S204: Obtain the set of interactive actions based on the category spacing between each interactive action.
[0103] After obtaining the category spacing between each interactive action, clustering can be performed on the category spacing. Then, the average category spacing in each category is calculated, and the average value is inverted.
[0104] In one optional example, a threshold can be preset to indicate that when the inverted average value is greater than the threshold, the similarity between interactive actions in the corresponding category is relatively high. For example, the category might contain both waving to the right and waving to the upper right. Such actions are easily confused by relevant devices, so one of these two actions should be deleted. This avoids having highly similar interactive actions in the final selection.
[0105] For ease of description, different Arabic numerals will be used below to distinguish different interactive actions. For example, interactive action 1 and interactive action 2 represent two different interactive actions.
[0106] For example, the class distance between interaction action 1 and interaction action 2 is 1, the class distance between interaction action 1 and interaction action 3 is 7, the class distance between interaction action 1 and interaction action 4 is 8, and the class distance between interaction action 3 and interaction action 4 is 2. After performing a clustering operation on these class distances, we can obtain the first class as {1 2} and the second class as {7 8}. Then, we calculate the average class distance in the first class as 1.5 and the average class distance in the second class as 7.5. Inverting these average values, for example, by adding a negative sign, yields -1.5 and -7.5.
[0107] In this example, the threshold can be set to -3. Since -1.5 is greater than -3 and -7.5 is less than -3, it indicates that the similarity between interactive actions in the first category is high, while the similarity between interactive actions in the second category is low. Therefore, even if the second category contains interactive actions 1, 3, and 4, the analysis of the first category shows that the similarity between interactive actions 3 and 4 is high. Thus, it is necessary to delete interactive actions 3 or 4 from the second category to ensure that the final selected interactive actions have a low degree of similarity.
[0108] In an optional example, when calculating the average of each category, the average of the first category and the average of the second category can be multiplied by a coefficient simultaneously, for example, by 2 to obtain an average of 3 and an average of 15. Then, a negative sign is added to the averages 3 and 15, and subsequent operations are performed. This disclosure does not limit the specific implementation process; that is, any modifications made based on the above solution that achieve the above objectives should be considered within the protection scope of this disclosure.
[0109] In one optional example, the false positive rate of each interactive action can be obtained first, and the false positive rates of each interactive action can be sorted. When the situation arises where it is necessary to delete interactive action 3 or interactive action 4 contained in the second category, the interactive action with the higher false positive rate can be selected for deletion.
[0110] For example, after sorting the false positive rates of each interactive action in ascending order, it is found that interactive action 3 is listed before interactive action 4, indicating that the false positive rate of interactive action 3 is lower than that of interactive action 4. Therefore, interactive action 4 in the second category can be deleted, resulting in interactive action 1 and interactive action 3. Interactive action 1 and interactive action 3 constitute the interactive action set.
[0111] Step S205: Obtain the usage parameter corresponding to each interactive action in the set of interactive actions. The usage parameter is used to characterize the user's usage preference for the interactive action.
[0112] Step S206: Obtain the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions. The recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized.
[0113] Step S207: For each interaction action, the usage parameter and the recognition parameter are fused to obtain the fused parameter corresponding to the interaction action.
[0114] Step S208: In response to the fusion parameters meeting the preset parameter selection conditions, the interactive action corresponding to the fusion parameters is determined as the target interactive action.
[0115] The interactive action selection method provided in this disclosure measures the similarity between different interactive actions by using the category gap, and pre-screens the interactive actions based on the category gap, effectively preventing the problem of confusion between devices and users caused by the high degree of similarity between the finally selected interactive actions.
[0116] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should know that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps may be performed in other orders or simultaneously.
[0117] Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by this disclosure.
[0118] Corresponding to the aforementioned application function implementation method embodiments, this disclosure also provides embodiments of application function implementation apparatus and corresponding terminals.
[0119] Figure 3 This is a schematic diagram of the structure of an interactive action selection device according to an exemplary embodiment of this disclosure, such as... Figure 3 As shown, the device for selecting the interactive action may include:
[0120] The first acquisition module 31 is used to acquire the usage parameters corresponding to each interactive action in the set of interactive actions, and the usage parameters are used to characterize the user's usage preference for the interactive action.
[0121] The second acquisition module 32 is used to acquire the recognition parameters corresponding to each interactive action in the set of interactive actions, and the recognition parameters are used to characterize the accuracy of the interactive action being recognized.
[0122] The fusion module 33 is used to fuse the usage parameter and the recognition parameter for each interaction action to obtain the fusion parameter corresponding to the interaction action.
[0123] The target interaction action determination module 34 is used to determine the interaction action corresponding to the fusion parameters as the target interaction action in response to the fusion parameters meeting the preset parameter selection conditions.
[0124] Optionally, the first acquisition module 31, when acquiring the usage parameters corresponding to each interactive action in the set of interactive actions, wherein the usage parameters are used to characterize the user's usage preference for the interactive action, includes:
[0125] For each interactive action, receive selection information corresponding to the interactive action, the selection information including at least one of the number of people who selected the interactive action and the frequency of use of the interactive action;
[0126] The selection information is statistically analyzed to obtain the usage parameter corresponding to the interactive action.
[0127] Optionally, the second acquisition module 32, when acquiring the recognition parameters corresponding to each interactive action in the set of interactive actions, wherein the recognition parameters are used to characterize the accuracy of the interactive action being recognized, includes:
[0128] For each interaction action, obtain the recognition rate and misjudgment rate corresponding to the interaction action;
[0129] The recognition rate and false positive rate of the interactive action are fused to obtain the recognition degree parameter corresponding to the interactive action. When fusing the recognition rate and the false positive rate, the higher the recognition rate, the higher the accuracy represented by the recognition degree parameter, and the lower the false positive rate, the higher the accuracy represented by the recognition degree parameter.
[0130] Optionally, the target interaction action determination module 34, when determining the interaction action corresponding to the fusion parameter as the target interaction action in response to the fusion parameter meeting preset parameter selection conditions, includes:
[0131] The fusion parameters corresponding to each interactive action in the set of interactive actions are sorted, and the interactive action corresponding to the fusion parameter ranked at a preset position is determined as the target interactive action based on the sorted fusion parameters.
[0132] Alternatively, in response to the value of the fusion parameter exceeding a preset parameter threshold, the interactive action corresponding to the fusion parameter is determined as the target interactive action.
[0133] Optional, such as Figure 4 As shown, the interactive action selection device, in addition to the following... Figure 3 In addition to the modules shown, it also includes:
[0134] The initial image acquisition module 41 is used to acquire an initial image set including multiple action images before selecting a portion of the action images as the target action from the set of action images to be selected, wherein each action image includes one action image.
[0135] The vector acquisition module 42 is used to extract features from each action image to obtain the feature vector corresponding to the interactive action.
[0136] The category spacing acquisition module 43 is used to obtain the category spacing between each interactive action based on the feature vectors corresponding to the interactive actions in each action image.
[0137] The interactive action set acquisition module 44 is used to acquire the interactive action set based on the category spacing between each interactive action.
[0138] Optionally, the vector acquisition module 42, when performing feature extraction on each action image to obtain the feature vector corresponding to the interaction action, includes:
[0139] The motion image is input into a category recognizer, wherein the category recognizer is used to identify the category of the interactive action in the motion image, and the category recognizer includes a feature extraction module and a category output module;
[0140] The feature extraction module in the category recognizer extracts features from the action image to obtain the feature vector corresponding to the interactive action. The category output module is used to obtain the category of the interactive action in the action image based on the feature vector output by the feature extraction module.
[0141] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0142] Accordingly, this disclosure provides an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to:
[0143] Obtain the usage parameter corresponding to each interactive action in the set of interactive actions, and the usage parameter is used to characterize the user's usage preference for the interactive action;
[0144] Obtain the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions, and the recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized;
[0145] For each of the interactive actions, the usage parameter and the recognition parameter are fused to obtain the fused parameter corresponding to the interactive action;
[0146] In response to the fusion parameters meeting preset parameter selection conditions, the interactive action corresponding to the fusion parameters is determined as the target interactive action.
[0147] Figure 5 This is a schematic diagram illustrating the structure of an electronic device 500 according to an exemplary embodiment. For example, the electronic device 500 can be a user device, specifically a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, wearable device such as smartwatch, smart glasses, smart bracelet, smart running shoes, etc.
[0148] Reference Figure 5 The electronic device 500 may include one or more of the following components: processing component 502, memory 504, power supply component 506, multimedia component 508, audio component 510, input / output (I / O) interface 512, sensor component 514, and communication component 516.
[0149] Processing component 502 typically controls the overall operation of electronic device 500, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 502 may include one or more processors 520 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 502 may include one or more modules to facilitate interaction between processing component 502 and other components. For example, processing component 502 may include a multimedia module to facilitate interaction between multimedia component 508 and processing component 502.
[0150] Memory 504 is configured to store various types of data to support the operation of device 500. Examples of this data include instructions for any application or method operating on device 500, contact data, phonebook data, messages, pictures, videos, etc. Memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0151] Power supply component 506 provides power to various components of device 500. Power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 500.
[0152] Multimedia component 508 includes a screen that provides an output interface between the aforementioned electronic device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0153] Audio component 510 is configured to output and / or input audio signals. For example, audio component 510 includes a microphone (MIC) configured to receive external audio signals when electronic device 500 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 504 or transmitted via communication component 516. In some embodiments, audio component 510 also includes a speaker for outputting audio signals.
[0154] I / O interface 512 provides an interface between processing component 502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0155] Sensor assembly 514 includes one or more sensors for providing state assessments of various aspects of electronic device 500. For example, sensor assembly 514 can detect the on / off state of electronic device 500, the relative positioning of components such as the display and keypad of electronic device 500, changes in position of electronic device 500 or a component of electronic device 500, the presence or absence of user contact with electronic device 500, orientation or acceleration / deceleration of electronic device 500, and temperature changes of electronic device 500. Sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 514 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0156] Communication component 516 is configured to facilitate wired or wireless communication between electronic device 500 and other devices. Electronic device 500 can access wireless networks based on communication standards, such as WiFi, 4G or 5G, 4G LTE, 5G NR, or combinations thereof. In one exemplary embodiment, communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the aforementioned communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0157] In an exemplary embodiment, the apparatus 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0158] In an exemplary embodiment, a non-transitory computer-readable storage medium, such as a memory 504 including instructions, is also provided. A selection method is provided to enable the electronic device 500 to perform an interactive action when the instructions in the storage medium are executed by the processor 520 of the electronic device 500. The method includes:
[0159] Obtain the usage parameter corresponding to each interactive action in the set of interactive actions, and the usage parameter is used to characterize the user's usage preference for the interactive action;
[0160] Obtain the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions, and the recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized;
[0161] For each of the interactive actions, the usage parameter and the recognition parameter are fused to obtain the fused parameter corresponding to the interactive action;
[0162] In response to the fusion parameters meeting preset parameter selection conditions, the interactive action corresponding to the fusion parameters is determined as the target interactive action.
[0163] The non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0164] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0165] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for selecting an interactive action, characterized in that, The method is used to select a subset of interactive actions as target interactive actions from a set of interactive actions to be selected. The method includes: For each interactive action in the set of interactive actions, receive selection information corresponding to the interactive action, the selection information including at least one of the number of people who selected the interactive action and the frequency of use of the interactive action; The selection information is statistically analyzed to obtain the usage parameter corresponding to the interactive action. The usage parameter is used to characterize the user's usage preference for the interactive action. Obtain the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions. The recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized. The recognition accuracy parameter corresponding to each interactive action is obtained by fusing the recognition rate and the false judgment rate corresponding to each interactive action. For each of the interactive actions, the usage parameter and the recognition parameter are fused to obtain the fused parameter corresponding to the interactive action; In response to the fusion parameters meeting preset parameter selection conditions, the interactive action corresponding to the fusion parameters is determined as the target interactive action.
2. The method according to claim 1, characterized in that, The step of obtaining the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions, wherein the recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized, includes: For each interaction action, obtain the recognition rate and misjudgment rate corresponding to the interaction action; The recognition rate and false positive rate of the interactive action are fused to obtain the recognition degree parameter corresponding to the interactive action. When fusing the recognition rate and the false positive rate, the higher the recognition rate, the higher the accuracy represented by the recognition degree parameter, and the lower the false positive rate, the higher the accuracy represented by the recognition degree parameter.
3. The method according to claim 1, characterized in that, The step of determining the interactive action corresponding to the fusion parameters as the target interactive action in response to the fusion parameters meeting preset parameter selection conditions includes: The fusion parameters corresponding to each interactive action in the set of interactive actions are sorted, and the interactive action corresponding to the fusion parameter ranked at a preset position is determined as the target interactive action based on the sorted fusion parameters. Alternatively, in response to the value of the fusion parameter exceeding a preset parameter threshold, the interactive action corresponding to the fusion parameter is determined as the target interactive action.
4. The method according to claim 1, characterized in that, Before obtaining the recognition parameters corresponding to each interactive action in the set of interactive actions, the method includes: Obtain an initial image set including multiple action images, wherein each action image includes an interactive action; Feature extraction is performed on each action image to obtain the feature vector corresponding to the interaction action; Based on the feature vectors corresponding to the interactive actions in each action image, the class distance between each interactive action is obtained; The set of interactive actions is obtained based on the category spacing between each interactive action.
5. The method according to claim 4, characterized in that, For each action image, feature extraction is performed to obtain the feature vector corresponding to the interaction action, including: The motion image is input into a category recognizer, wherein the category recognizer is used to identify the category of the interactive action in the motion image, and the category recognizer includes a feature extraction module and a category output module; The feature extraction module in the category recognizer extracts features from the action image to obtain the feature vector corresponding to the interactive action. The category output module is used to obtain the category of the interactive action in the action image based on the feature vector output by the feature extraction module.
6. An interactive action selection device, characterized in that, The device is used to select a subset of interactive actions as target interactive actions from a set of interactive actions to be selected. The device includes: The first acquisition module is configured to receive selection information corresponding to each interaction action in the set of interaction actions, wherein the selection information includes at least one of the number of people who selected the interaction action and the frequency of use of the interaction action; and to perform statistics on the selection information to obtain a usage parameter corresponding to the interaction action, wherein the usage parameter is used to characterize the user's usage preference for the interaction action. The second acquisition module is used to acquire the recognition parameters corresponding to each interactive action in the set of interactive actions. The recognition parameters are used to characterize the accuracy of the interactive action being recognized. The recognition parameters corresponding to each interactive action are obtained by fusing the recognition rate and the false judgment rate corresponding to each interactive action. The fusion module is used to fuse the usage parameter and the recognition parameter for each interaction action to obtain the fusion parameter corresponding to the interaction action. The target interaction action determination module is used to determine the interaction action corresponding to the fusion parameters as the target interaction action in response to the fusion parameters meeting the preset parameter selection conditions.
7. The apparatus according to claim 6, characterized in that, The second acquisition module, when acquiring the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions, wherein the recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized, includes: For each interaction action, obtain the recognition rate and misjudgment rate corresponding to the interaction action; The recognition rate and false positive rate of the interactive action are fused to obtain the recognition degree parameter corresponding to the interactive action. When fusing the recognition rate and the false positive rate, the higher the recognition rate, the higher the accuracy represented by the recognition degree parameter, and the lower the false positive rate, the higher the accuracy represented by the recognition degree parameter.
8. The apparatus according to claim 6, characterized in that, The target interaction action determination module, when determining the interaction action corresponding to the fusion parameters as the target interaction action in response to the fusion parameters meeting preset parameter selection conditions, includes: The fusion parameters corresponding to each interactive action in the set of interactive actions are sorted, and the interactive action corresponding to the fusion parameter ranked at a preset position is determined as the target interactive action based on the sorted fusion parameters. Alternatively, in response to the value of the fusion parameter exceeding a preset parameter threshold, the interactive action corresponding to the fusion parameter is determined as the target interactive action.
9. The apparatus according to claim 6, characterized in that, The device includes: An initial image acquisition module is used to acquire an initial image set including multiple action images before selecting a portion of the action images as target action images from the set of action images to be selected, wherein each action image includes one action image. The vector acquisition module is used to extract features from each action image to obtain the feature vector corresponding to the interaction action; The category spacing acquisition module is used to obtain the category spacing between interactive actions based on the feature vectors corresponding to the interactive actions in each action image. The interactive action set acquisition module is used to acquire the interactive action set based on the category spacing between each interactive action.
10. The apparatus according to claim 9, characterized in that, The vector acquisition module, when performing feature extraction on each action image to obtain the feature vector corresponding to the interaction action, includes: The motion image is input into a category recognizer, wherein the category recognizer is used to identify the category of the interactive action in the motion image, and the category recognizer includes a feature extraction module and a category output module; The feature extraction module in the category recognizer extracts features from the action image to obtain the feature vector corresponding to the interactive action. The category output module is used to obtain the category of the interactive action in the action image based on the feature vector output by the feature extraction module.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1 to 5.
12. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store processor-executable instructions; The processor is configured as follows: For each interactive action in the set of interactive actions, receive selection information corresponding to the interactive action, the selection information including at least one of the number of people who selected the interactive action and the frequency of use of the interactive action; The selection information is statistically analyzed to obtain the usage parameter corresponding to the interactive action. The usage parameter is used to characterize the user's usage preference for the interactive action. Obtain the recognition accuracy parameter corresponding to each interactive action in the set of interactive actions. The recognition accuracy parameter is used to characterize the accuracy of the interactive action being recognized. The recognition accuracy parameter corresponding to each interactive action is obtained by fusing the recognition rate and the false judgment rate corresponding to each interactive action. For each of the interactive actions, the usage parameter and the recognition parameter are fused to obtain the fused parameter corresponding to the interactive action; In response to the fusion parameters meeting preset parameter selection conditions, the interactive action corresponding to the fusion parameters is determined as the target interactive action.
Citation Information
Patent Citations
Human body action hierarchical identification method based on pyroelectric infrared detection
CN102567715A
Method for acquiring user operation habit information, diagnosis device and server
CN110298935A