User behavior recognition method by screen recorded data, and device and readable storage medium
By extracting and classifying key frame images from screen recordings, the method enhances user behavior analysis efficiency and accuracy, addressing the inefficiencies and inaccuracies of manual methods.
Patent Information
- Application Number
- JP2024109335
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2024-07-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing user behavior analysis methods based on screen recording data are inefficient and inaccurate due to high human resource costs and varying quality, affecting the accuracy of analysis results.
Extract key frame images from screen recording data, perform data analysis to extract feature information, and classify these images to obtain classification information, which is used to characterize user operations, and traverse the classification information of consecutive frames to recognize user behavior.
Improves the efficiency and accuracy of user behavior recognition by automating the process and reducing misrecognition rates.
Smart Images

Figure 2025106182000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of user behavior recognition, and particularly to a method and apparatus for user behavior recognition based on screen recording data, and a readable storage medium.
Background Art
[0002] With the development of mobile terminal and network technologies, the screen recording technology is used to record the video of the application program operation interface, and combined with methods such as manual detection, to determine the operations performed by the user in the screen recording, and to obtain the user's operation information from the recorded video. Currently, generally, the action data including brands, products, advertisements, etc. included in the video stream is selected one by one in the way of manual search and labeling, and the user behavior analysis is performed based on the selected action data. However, such a method of performing user behavior analysis manually requires a lot of human costs and is also inefficient. In addition, since the quality of manual analysis varies, it affects the accuracy of the user behavior analysis results.
Summary of the Invention
Problems to be Solved by the Invention
[0003] The technical problem to be solved by the embodiments of the present invention is how to perform user behavior analysis efficiently and accurately.
Means for Solving the Problems
[0004] To solve the above technical problems, an embodiment of the present invention provides a method for recognizing user behavior based on screen recording data, including the steps of extracting key frame images from a plurality of frame images in the screen recording data to obtain key frame images of a plurality of frames; performing data analysis on the key frame images of each frame to extract feature information of the key frame images of each frame; performing screen classification on the key frame images according to the feature information of the key frame images of each frame to obtain classification information of the key frame images of each frame, where the classification information is used to characterize the operation actions performed by the user in the key frame images; traversing the classification information of the key frame images of each frame in the screen recording data, and obtaining a user behavior recognition result based on the correlation between the classification information of the key frame images of a plurality of consecutive frames.
[0005] Preferably, the step of extracting key frame images from a plurality of frame images in the screen recording data includes comparing the screen information changes of adjacent images in the screen recording data, and when the proportionality of the screen information changes is greater than a predetermined change threshold, setting the subsequent frame image of the adjacent images as the key frame image.
[0006] Preferably, after obtaining the key frame images, it further includes determining whether the time interval between adjacent key frame images is greater than a preset shortest time interval, and if it is not greater than the preset shortest time interval, deleting the key frame image of the subsequent frame of the adjacent key frame images.
[0007] Preferably, the step of performing data analysis on the key frame image of each frame and extracting the feature information of the key frame image of each frame includes: the feature information includes character feature information, and performing optical character recognition on the key frame image of each frame to extract the character feature information in the key frame image of each frame; and / or the feature information includes target feature information, and performing target detection on the key frame image of each frame to extract the target feature information in the key frame image of each frame.
[0008] Preferably, the user behavior recognition method further includes the step of calculating the amount of displacement of the position of the character feature information in adjacent key frame images for adjacent key frame images, and deleting the key frame image of the later frame among the adjacent key frame images when the amount of displacement is smaller than a predetermined displacement threshold.
[0009] Preferably, the step of obtaining the user behavior recognition result based on the correlation between the classification information of the key frame images of a plurality of consecutive frames includes: the correlation includes the dependency relationship in the time dimension of the user operation behavior characterized by the classification information, and determining the user behavior recognition result based on the dependency relationship in the time dimension of the user operation behavior characterized by the classification information of the key frame images of a plurality of consecutive frames.
[0010] Preferably, a dependency relationship in the time dimension of user operation actions is characterized by a predetermined action sequence. The predetermined action sequence includes a plurality of classification information indicating a plurality of continuous operation actions corresponding to user actions. Based on the dependency relationship in the time dimension of user operation actions characterized by the classification information of key frame images of a plurality of continuous frames, the step of determining the user action recognition result includes: determining whether the classification information of key frame images of a plurality of continuous frames matches the predetermined action sequence; and when the classification information of key frame images of a plurality of continuous frames matches the predetermined action sequence, obtaining the user action recognition result based on the user action characterized by the predetermined action sequence.
[0011] Preferably, the step of determining whether the classification information of the key frame images of a plurality of consecutive frames matches the predetermined action sequence includes: obtaining a predetermined action sequence associated with the classification information of the key frame image of the traversed i-th frame, where i is a positive integer greater than or equal to 1; when traversing to the key frame image of the (i + 1)-th frame, obtaining a predetermined action sequence associated with the classification information of the key frame image of the (i + 1)-th frame, and determining whether the classification information of the key frame image of the (i + 1)-th frame matches the predetermined action sequence associated with the key frame image of the i-th frame; when the classification information of the key frame image of the (i + 1)-th frame matches the predetermined action sequence associated with the key frame image of the i-th frame, retaining the predetermined action sequence associated with the key frame image of the i-th frame, and continuing to match the classification information of the key frame image of the (i + 2)-th frame with the predetermined action sequence associated with the key frame image of the i-th frame until the classification information of the key frame images of a plurality of consecutive frames matches the predetermined action sequence associated with the key frame image of the i-th frame and the user action recognition result is obtained; when the classification information of the key frame image of the (i + 1)-th frame does not match the predetermined action sequence associated with the key frame image of the i-th frame, discarding the predetermined action sequence that does not match.
[0012] Preferably, the step of obtaining a predetermined action sequence associated with the classification information of the key frame image of the traversed i-th frame includes: matching the classification information of the key frame image of the i-th frame with the first classification information in each predetermined action sequence, and setting the matched predetermined action sequence as the predetermined action sequence associated with the classification information of the key frame image of the i-th frame.
[0013] Preferably, before traversing the classification information of the key frame images of each frame in the screen recording data, the method further includes screening the key frame images based on the classification information, and using the screened key frame images as the key frame images to be traversed.
[0014] Preferably, the step of screening the key frame images based on the classification information includes: determining whether there are continuous multiple frames of key frame images having the same classification information according to the time series order of the key frame images; and when there are continuous multiple frames of key frame images having the same classification information, retaining one of the continuous multiple frames of key frame images having the same classification information.
[0015] Preferably, the user behavior recognition result includes the duration of the user behavior and / or auxiliary information of the user behavior. The duration of the user behavior is obtained by taking the total duration of the continuous multiple frames of key frame images as the duration of the user behavior. The auxiliary information of the user behavior is obtained by identifying candidate key frame images from the continuous multiple frames of key frame images according to the classification information of the key frame images of each frame based on the operation actions to be the target of obtaining auxiliary information in the predetermined action sequence, and obtaining the auxiliary information based on the feature information of the candidate key frame images.
[0016] Preferably, the predetermined action sequence is obtained by acquiring the application identifier corresponding to the screen recording data from the log information of the screen recording data, and acquiring the predetermined action sequence associated with the application identifier.
[0017] Embodiments of the present invention further provide a user behavior recognition device based on screen recording data, including a key frame image extraction unit that extracts key frame images from multiple frame images in the screen recording data to obtain key frame images of multiple frames, a feature information extraction unit that performs data analysis on the key frame images of each frame to extract feature information of the key frame images of each frame, and a classification unit that performs screen classification on the key frame images according to the feature information of the key frame images of each frame to obtain classification information of the key frame images of each frame, where the classification information is used to characterize the operation actions performed by the user in the key frame images, and a user behavior recognition unit that traverses the classification information of the key frame images of each frame in the screen recording data and obtains a user behavior recognition result based on the correlation between the classification information of the key frame images of multiple consecutive frames.
[0018] Embodiments of the present invention further provide a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, it performs the steps of any of the above user behavior recognition methods based on screen recording data.
[0019] Embodiments of the present invention further provide a user behavior recognition device including a processor and a memory storing a computer program executable by the processor, where when the processor executes the computer program, it performs the steps of any of the above user behavior recognition methods based on screen recording data.
Advantages of the Invention
[0020] Compared with the prior art, the technical solutions of the embodiments of the present invention have the following beneficial effects. By performing key frame extraction on multiple frame images in the screen recording data, key frame images of multiple frames are obtained, and feature information of the key frame images of multiple frames is extracted. Screen classification is performed on the key frame images according to the feature information of the key frame images of each frame to obtain classification information of the key frame images of each frame. Since the classification information can be used to characterize the operation actions performed by the user in the key frame images, the user behavior recognition result can be obtained based on the relationship between the classification information of the key frame images of multiple consecutive frames. According to the above method, by obtaining the user behavior recognition result based on the screen recording data, the efficiency and accuracy of user behavior recognition can be improved.
Brief Description of the Drawings
[0021]
Figure 1
Figure 2
Modes for Carrying Out the Invention
[0022] As described above, in the prior art, the traditional method of manually recognizing user behavior is costly in terms of human resources and inefficient. In addition, since there are variations in the quality of the analysis by manual work, it affects the accuracy of the user behavior analysis result.
[0023] In order to solve the above problems, in an embodiment of the present invention, key frame extraction is performed on a plurality of frame images in screen recording data to obtain key frame images of a plurality of frames, and feature information of the key frame images of the plurality of frames is extracted. Screen classification is performed on the key frame images according to the feature information of the key frame images of each frame to obtain classification information of the key frame images of each frame. Since the classification information can be used to characterize the operation actions performed by the user in the key frame images, a user behavior recognition result can be obtained based on the correlation between the classification information of the key frame images of a plurality of consecutive frames. According to the above method, by obtaining a user behavior recognition result based on screen recording data, the efficiency and accuracy of user behavior recognition can be improved.
[0024] In order to make the above objects, features and beneficial effects of the embodiments of the present invention clearer, the following will refer to the drawings to describe the specific embodiments of the present invention in detail.
[0025] The embodiments of the present invention provide a method for recognizing user behavior based on screen recording data. The method for recognizing user behavior may be executed by a terminal device, or may be executed by a server or a cloud platform, etc. The terminal device may include a suitable terminal such as a computer or a notebook computer.
[0026] Referring to FIG. 1, the method for recognizing user behavior based on screen recording data specifically includes the following steps 11 to 14.
[0027] Step 11: Perform key frame image extraction on a plurality of frame images in the screen recording data to obtain key frame images of a plurality of frames.
[0028] Step 12: Perform data analysis on the key frame images of each frame to extract the feature information of the key frame images of each frame.
[0029] Step 13: Classify the key-frame images of each frame according to the feature information of the key-frame images of each frame, obtain the classification information of the key-frame images of each frame, and the classification information is used to characterize the operation actions performed by the user in the key-frame images.
[0030] Step 14: Traverse the classification information of the key-frame images of each frame in the screen recording data, and obtain the user behavior recognition result based on the correlation between the classification information of the key-frame images of multiple consecutive frames.
[0031] As can be seen from the above, by performing key-frame extraction on multiple frame images in the screen recording data, key-frame images of multiple frames are obtained, and the feature information of the key-frame images of multiple frames is extracted. Classify the key-frame images according to the feature information of the key-frame images of each frame, obtain the classification information of the key-frame images of each frame, and the classification information can be used to characterize the operation actions performed by the user in the key-frame images. Therefore, based on the correlation between the classification information of the key-frame images of multiple consecutive frames, the user behavior recognition result can be obtained. According to the above method, by obtaining the user behavior recognition result based on the screen recording data, the efficiency and accuracy of user behavior recognition can be improved.
[0032] In addition, by adopting the above-mentioned method, the recognition of a series of operation actions of the user can be realized. Furthermore, compared with the conventional method in which only a single operation action can be recognized, the accuracy of user behavior recognition can be further improved, and the misrecognition rate of user behavior can be reduced.
[0033] In a specific implementation, one or more application software can be installed on the terminal device. After obtaining the authorization and permission from the user, when the user operates the application software on the terminal device, video recording is performed on the screen displayed on the user interface of the terminal device to obtain screen recording data.
[0034] Screen recording data can be divided into screens to obtain continuous images of multiple frames.
[0035] In a specific implementation of step 11, key frame images can be extracted from multiple frame images in the screen recording data as follows. Specifically, the change in screen information of adjacent images in the screen recording data is compared. When the proportionality of the change in screen information is greater than a predetermined change threshold, the subsequent frame image among the adjacent images is used as the key frame image. The predetermined change threshold can be set based on factors such as the processing capacity of the terminal device and the tolerance for information loss. The higher the predetermined change threshold, the greater the change in screen information between adjacent key frame images, and the greater the loss of all the key frame images extracted from the screen information included in the screen recording data. At this time, the number of obtained key frame images is relatively small, the requirement for the processing capacity of the terminal is low, and the tolerance for information loss is high. Correspondingly, the lower the predetermined change threshold, the relatively smaller the change in screen information between adjacent key frame images, and the relatively smaller the loss of all the key frame images extracted from the screen information included in the screen recording data.
[0036] Generally, when extracting key frame images based on the change in screen information, the tolerance for information loss is small, that is, the similarity between adjacent key frame images is relatively large. When recognizing user behavior, generally, since user operation actions continue for a certain period of time, multiple effective operation actions are not generated within a short period of time.
[0037] In order to improve the effectiveness of user behavior recognition in the extracted key-frame images and the extraction efficiency of the key-frame images, in some embodiments of the present invention, after extracting and obtaining the key-frame images, the key-frame images may be further screened. For example, it is determined whether the time interval between adjacent key-frame images is greater than a preset minimum time interval. If it is not greater than the preset minimum time interval, the key-frame image of the later frame among the adjacent key-frame images is deleted.
[0038] In some embodiments, an application identifier corresponding to the screen recording data may be obtained, and the preset minimum time interval may be determined based on the application identifier. The application identifier is used to characterize the application software corresponding to the screen recording data, and different application identifiers may each have a corresponding preset minimum time interval, and the preset minimum time intervals corresponding to different application identifiers may be different.
[0039] In some non-limiting embodiments, the application identifier corresponding to the screen recording data can be obtained from the log information of the screen recording data.
[0040] In a specific implementation, after extracting and obtaining key-frame images of multiple frames, based on the positions of the key-frame images of multiple frames in the screen recording data, in order to identify the relative positional order of the key-frame images of multiple frames, the key-frame images of multiple frames can be marked. The marking may be to number the key-frame images of multiple frames. Of course, the time of the key-frame images in the screen recording data may also be used as marking information.
[0041] In the specific implementation of step 12, the feature information of the key-frame image of each frame may include at least one of character feature information and target feature information.
[0042] In some embodiments, optical character recognition (OCR) is performed on the key-frame images of each frame to extract the character feature information in the key-frame images of each frame.
[0043] In a specific implementation, the character feature information includes character content information, position information of the character content in the key-frame image, and the like. For adjacent key-frame images, the amount of displacement of the character feature information in the adjacent key-frame images is calculated. If the amount of displacement is smaller than a predetermined displacement threshold, the key-frame image of the subsequent frame among the adjacent key-frame images is deleted. The character feature information can characterize the information related to the user behavior carried by the key-frame image. When the amount of displacement of the character feature information in the adjacent key-frame images is small, it can indicate that the difference between the key-frame images of the adjacent frames is small. In this way, by deleting the key-frame images of the subsequent frames, the accuracy of user behavior recognition can be ensured, the simplification of the key-frame images can be realized, and the user behavior efficiency can be improved.
[0044] For example, the character content information can include one or more of product value, product name, names of each navigation bar included in the site navigation page, name of the button, and the like. It should be noted that the specific content included in the character content information varies depending on different application scenarios and is not limited here.
[0045] In some non-limiting embodiments, for adjacent key-frame images, it can be checked whether the character content information in the character feature information in the key-frame images of two adjacent frames is the same. If the character content information is the same, the amount of displacement of the same character feature information is calculated. If the amount of displacement is smaller than a predetermined displacement threshold, the key-frame image of the subsequent frame among the adjacent key-frame images is deleted.
[0046] In some other embodiments, target detection is performed on the key-frame images of each frame to extract target feature information in the key-frame images of each frame. The target feature information may be brand information, product information, etc. The product information may include product pictures.
[0047] In some other embodiments, the amount of displacement of the character feature information in adjacent key-frame images can be calculated. When the amount of displacement is smaller than a predetermined displacement threshold, the key-frame image of the subsequent frame among the adjacent key-frame images is deleted. After screening the key-frame images of multiple frames, target detection is performed on the key-frame images of the multiple frames after screening to extract target feature information in the key-frame images of each frame.
[0048] In some embodiments, after extracting the feature information of the key-frame image of each frame, the relationship between the key-frame image and the feature information can be established, and the feature information of the key-frame image and the relationship between the key-frame image and the feature information can be memorized.
[0049] In a specific implementation, the classification information is used to characterize the operation actions performed by the user in the key-frame image. If the application scenarios are different, the application software corresponding to the screen recording data is different, and the operation actions performed by the user in the key-frame image are different. Taking shopping applications such as Taobao and JD.com as examples, the operation actions performed by the user in the key-frame image can include product browsing, entering the product details page, adding to the shopping cart, placing an order, making a payment, etc. Each classification information corresponds to one or more pieces of feature information respectively, and the relationship between the classification information and the feature information may be set. Then, based on the feature information such as the key-frame image, characters, and positions of each frame, screen classification can be performed on the key-frame image.
[0050] The classification information of the key frame images of each frame may be one or more.
[0051] The related relationship includes a dependency relationship in the time dimension of the user operation actions characterized by the classification information. The dependency relationship in the time dimension can represent that the user operation actions characterized by the classification information of adjacent key frames have a front-back dependency relationship in the time dimension. For example, the user operation actions characterized by the classification information of the key frame images of a later frame depend on the user operation actions characterized by the classification information of the key frame images of a previous frame.
[0052] In some embodiments, a predetermined action sequence characterizes the dependency relationship in the time dimension of the user operation actions, and the predetermined action sequence includes a plurality of classification information indicating a plurality of consecutive operation actions corresponding to the user actions.
[0053] A plurality of predetermined action sequences can be set in advance. The predetermined action sequence is obtained as follows: obtain the application identifier of the screen recording data from the log information of the screen recording data, and obtain the predetermined action sequence associated with the application identifier. The application identifier is for identifying application software and corresponds one by one to the application software. The application identifier may be the number, name, etc. of the application software.
[0054] In the specific implementation of step 14, it is determined whether the classification information of the key frame images of a plurality of consecutive frames matches the predetermined action sequence. If the classification information of the key frame images of a plurality of consecutive frames matches the predetermined action sequence, the user action recognition result is obtained based on the user actions characterized by the predetermined action sequence.
[0055] In a specific implementation, each user action can correspond to one or more predetermined action sequences. When performing user action recognition, the obtained predetermined action sequence may be a predetermined action sequence of one user action or a predetermined action sequence of multiple user actions. Regarding each predetermined action sequence as a link, when performing user action recognition, the classification information of the key frame image of each frame may be matched one by one with the predetermined action sequences of each link, or the classification information of the key frame image of each frame may be matched simultaneously with the predetermined action sequences of all links.
[0056] In some non-limiting embodiments, when traversing the classification information of the key frame image of each frame in the screen recording data, a predetermined action sequence associated with the classification information of the key frame image of the i-th frame traversed is obtained. Here, i is a positive integer greater than or equal to 1. When traversing to the key frame image of the (i + 1)-th frame, a predetermined action sequence associated with the classification information of the key frame image of the (i + 1)-th frame is obtained, and it is determined whether the classification information of the key frame image of the (i + 1)-th frame matches the predetermined action sequence associated with the key frame image of the i-th frame. If the classification information of the key frame image of the (i + 1)-th frame matches the predetermined action sequence associated with the key frame image of the i-th frame, the predetermined action sequence associated with the key frame image of the i-th frame is retained, and the matching between the classification information of the key frame images of a plurality of consecutive frames and the predetermined action sequence associated with the key frame image of the i-th frame is continued until the user action recognition result is obtained. If the classification information of the key frame image of the (i + 1)-th frame does not match the predetermined action sequence associated with the key frame image of the i-th frame, the non-matching predetermined action sequence is discarded.
[0057] Also, when traversing to the key frame image of the i+1-th frame, a predetermined action sequence associated with the classification information of the key frame image of the i+1-th frame is obtained. When traversing to the key frame image of the i+2-th frame, it is determined whether the classification information of the key frame image of the i+2-th frame matches the predetermined action sequence associated with the key frame image of the i+1-th frame. If the classification information of the key frame image of the i+2-th frame matches the predetermined action sequence associated with the key frame image of the i+1-th frame, the predetermined action sequence associated with the key frame image of the i+1-th frame is retained, and until the classification information of the key frame images of a plurality of consecutive frames matches the predetermined action sequence and the user action recognition result is obtained, the classification information of the key frame images after the i+2-th frame is matched with the predetermined action sequence associated with the key frame image of the i+1-th frame. Or, if the classification information of the key frame image of the i+2-th frame does not match the predetermined action sequence associated with the key frame image of the i+1-th frame, the predetermined action sequence associated with the key frame image of the i+1-th frame that does not match the classification information of the key frame image of the i+2-th frame is discarded.
[0058] In a specific implementation, the classification information of the key frame image of the i-th frame is matched with the first classification information in each predetermined action sequence, and the matched predetermined action sequence is used as the predetermined action sequence associated with the classification information of the key frame image of the i-th frame.
[0059] If there is a plurality of classification information of the key frame image of the i-th frame, each classification information is respectively matched with the first classification information in each predetermined action sequence, and the matched predetermined action sequence is used as the predetermined action sequence associated with the classification information of the key frame image of the i-th frame. The predetermined action sequence associated with the classification information of the key frame image of the i-th frame may be one or a plurality.
[0060] In some non-limiting embodiments, before traversing the classification information of the key-frame image of each frame in the screen recording data, the key-frame images are further screened based on the classification information, and the screened key-frame images are used as the key-frame images to be traversed.
[0061] In some embodiments, the key-frame images may be screened based on the classification information as follows. Specifically, it is determined whether there are key-frame images of a plurality of consecutive frames having the same classification information according to the time series order of the key-frame images. If there are key-frame images of a plurality of consecutive frames having the same classification information, one key-frame image of the key-frame images of the plurality of consecutive frames having the same classification information is reserved. For example, the first key-frame image of the key-frame images of the plurality of consecutive frames may be reserved, or the last key-frame image may be reserved, or any one key-frame image of the key-frame images of the plurality of consecutive frames may be reserved. After screening the key-frame images based on the classification information, the obtained key-frame images of the plurality of frames are used as the key-frame images to be traversed. Thereby, by screening the key-frame images according to the classification information and traversing based on the screened key-frame images, it is possible to stay on the key-frame images having the same classification information and repeatedly match against the same predetermined action sequence, or avoid matching against some pages not related to the classification information (for example, blank pages, or staying on the same key-frame image due to network delay, etc.), and improve the matching efficiency.
[0062] In some embodiments, when there are key-frame images having a plurality of classification information among the key-frame images of a plurality of consecutive frames having the same classification information, the key-frame images having the plurality of classification information are reserved.
[0063] In a specific implementation, the user behavior recognition result further includes the duration of the user behavior and / or auxiliary information of the user behavior.
[0064] In some embodiments, the duration of the user behavior is obtained by taking the sum of the durations of the key frame images of the plurality of consecutive frames as the duration of the user behavior.
[0065] In some other embodiments, the auxiliary information of the user behavior is obtained as follows: based on the operation actions to be the target of obtaining auxiliary information in the predetermined behavior sequence, candidate key frame images are identified from the key frame images of the plurality of frames according to the classification information of the key frame images of each frame, and the auxiliary information is obtained based on the feature information of the candidate key frame images.
[0066] The auxiliary information includes information on related objects of the user behavior. For example, when the user behavior is an order placement, the information on the related object may be the name of the ordered product, the quantity of the product, or the amount of the product. Also, for example, when the user behavior is page browsing, the information on the related object may be the name of the browsed page, the browsing time, or the browsing portal.
[0067] Furthermore, the relationship between the acquisition rule of the auxiliary information and the user behavior can be arranged. The acquisition rule of the auxiliary information is used to indicate the classification information where the auxiliary information appears. After determining the user behavior in the screen recording data, the acquisition rule of the auxiliary information is obtained, candidate key frame images are determined based on the acquisition rule of the auxiliary information and the classification information of the key frame images of each frame, and the auxiliary information is obtained from the feature information of the candidate key frame images.
[0068] In a specific implementation, the user behavior recognition result can be saved as a dataset.
[0069] Referring to FIG. 2, an embodiment of the present invention further provides a user behavior recognition device based on screen recording data. The user behavior recognition device 20 includes a key frame image extraction unit 21 that extracts key frame images from a plurality of frame images in the screen recording data to obtain key frame images of a plurality of frames, a feature information extraction unit 22 that performs data analysis on the key frame images of each frame to extract feature information of the key frame images of each frame, and a classification unit 23 that performs screen classification on the key frame images according to the feature information of the key frame images of each frame to obtain classification information of the key frame images of each frame, where the classification information is used to characterize the operation actions performed by the user in the key frame images, and a user behavior recognition unit 24 that traverses the classification information of the key frame images of each frame in the screen recording data and obtains a user behavior recognition result based on the correlation between the classification information of the key frame images of a plurality of consecutive frames.
[0070] In a specific implementation, the user behavior recognition device 20 is used to implement the above user behavior recognition method. The user behavior recognition device 20 includes units for implementing each step in the user behavior recognition method. For the specific operation principle and operation flow of the user behavior recognition device 20, reference may be made to the description of the user behavior recognition method in the above embodiment, and details will not be repeated here.
[0071] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it executes the steps of the user behavior recognition method based on screen recording data provided by any of the above embodiments of the present invention.
[0072] The computer-readable storage medium may include non-volatile memory or non-transitory memory, or may include optical disks, mechanical hard disks, solid-state drives, etc.
[0073] An embodiment of the present invention further provides a user behavior recognition device based on screen recording data, including a processor and a memory storing a computer program executable by the processor. When the processor executes the computer program, it performs the steps of the user behavior recognition method based on screen recording data provided by any of the above embodiments.
[0074] The memory and the processor may be combined and located inside the user behavior recognition device based on screen recording data, or may be located outside the user behavior recognition device based on screen recording data. The memory and the processor may be connected via a communication bus.
[0075] The user behavior recognition device based on screen recording data may include terminal devices such as mobile phones, computers, tablets, etc., but is not limited thereto, and may be a server, a cloud platform, etc.
[0076] The above embodiments may be implemented in whole or in part by software, hardware, firmware, or any other arbitrary combination. When implemented by software, the above embodiments may be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, they generate all or part of the flow or functions described in the embodiments of the present application. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer program may be stored in a computer-readable storage medium, or may be transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, it may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner.
[0077] In some embodiments provided by the present application, the disclosed methods, apparatuses, and systems may be implemented in other forms. For example, the apparatus embodiments described above are merely schematic. For example, the partitioning of the above units is merely a partitioning of logical functions. In actual implementation, there may be other partitioning methods. For example, a plurality of units or components may be combined, integrated into another system, or some features may be ignored or not executed. The units described as individual members may or may not be physically separated. The members shown as units may or may not be physical units, that is, they may be located in a single place or distributed among a plurality of network units. Depending on actual needs, some or all of these units may be selected to achieve the purpose of the solution of this embodiment.
[0078] Also, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may be physically included individually, or two or more units may be integrated into one unit. The above integrated units may be implemented in the form of hardware or in the form of hardware + software functional units.
[0079] Here, the term "and / or" in this specification merely describes the relationship of related objects and indicates that three relationships exist. For example, A and / or B indicates three situations: A exists individually, A and B exist simultaneously, and B exists individually. Also, the symbol " / " in this specification indicates that the objects related before and after are in an "or" relationship.
[0080] "Plurality" in the embodiments of the present application refers to two or more.
[0081] Here, the numbers of each step in this embodiment do not limit the execution order of each step.
[0082] Although the present invention has been disclosed as above, it is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope limited by the claims.
Claims
1. A method for recognizing user behavior based on screen recording data, comprising: extracting key frame images from a plurality of frame images in the screen recording data to obtain key frame images of the plurality of frames; performing data analysis on the key frame images of each frame to extract feature information of the key frame images of each frame; performing screen classification on the key frame images according to the feature information of the key frame images of each frame to obtain classification information of the key frame images of each frame, wherein the classification information is used to characterize the operation actions performed by the user in the key frame images; traversing the classification information of the key frame images of each frame in the screen recording data, and obtaining a user behavior recognition result based on the correlation between the classification information of the key frame images of a plurality of consecutive frames; A user behavior recognition method characterized by comprising the above steps.
2. The step of extracting key frame images from a plurality of frame images in the screen recording data is: comparing the screen information changes of adjacent images in the screen recording data, and when the proportionality of the screen information changes is greater than a predetermined change threshold, setting the subsequent frame image of the adjacent images as the key frame image. The user behavior recognition method according to claim 1, characterized by comprising the above steps.
3. After obtaining the key frame images, further: determining whether the time interval between adjacent key frame images is greater than a preset minimum time interval, and if it is not greater than the preset minimum time interval, deleting the key frame image of the subsequent frame of the adjacent key frame images. The user behavior recognition method according to claim 1, characterized by comprising the above steps.
4. The step of performing data analysis on the key frame images of each frame to extract feature information of the key frame images of each frame is: the feature information includes character feature information, and performing optical character recognition on the key frame images of each frame to extract the character feature information in the key frame images of each frame, and / or the feature information includes target feature information, and performing target detection on the key frame images of each frame to extract the target feature information in the key frame images of each frame. The user behavior recognition method according to claim 1, characterized by including
5. For adjacent key-frame images, calculating the amount of displacement of the position of the character feature information in the adjacent key-frame images, and when the amount of displacement is smaller than a predetermined displacement amount threshold, deleting the key-frame image of the later frame among the adjacent key-frame images The user behavior recognition method according to claim 4, further characterized by further including
6. Based on the correlation between the classification information of the key-frame images of a plurality of consecutive frames, the step of obtaining the user behavior recognition result is The correlation includes a dependency relationship in the time dimension of the user operation action characterized by the classification information, and based on the dependency relationship in the time dimension of the user operation action characterized by the classification information of the key-frame images of a plurality of consecutive frames, the step of determining the user behavior recognition result The user behavior recognition method according to claim 1, characterized by including
7. Characterize the dependency relationship in the time dimension of the user operation action by a predetermined action sequence, and the predetermined action sequence includes a plurality of classification information indicating a plurality of consecutive operation actions corresponding to the user behavior Based on the dependency relationship in the time dimension of the user operation action characterized by the classification information of the key-frame images of a plurality of consecutive frames, the step of determining the user behavior recognition result is The step of determining whether the classification information of the key-frame images of a plurality of consecutive frames matches the predetermined action sequence When the classification information of the key-frame images of a plurality of consecutive frames matches the predetermined action sequence, the step of obtaining the user behavior recognition result based on the user behavior characterized by the predetermined action sequence The user behavior recognition method according to claim 6, characterized by including
8. The step of determining whether the classification information of the key-frame images of a plurality of consecutive frames matches the predetermined action sequence is The step of obtaining a predetermined action sequence associated with the classification information of the i-th key-frame image traversed, where i is a positive integer greater than or equal to 1 When traversing to the key frame image of the (i + 1)-th frame, obtain a predetermined action sequence associated with the classification information of the key frame image of the (i + 1)-th frame, and determine whether the classification information of the key frame image of the (i + 1)-th frame matches the predetermined action sequence associated with the key frame image of the i-th frame; When the classification information of the key frame image of the (i + 1)-th frame matches the predetermined action sequence associated with the key frame image of the i-th frame, retain the predetermined action sequence associated with the key frame image of the i-th frame, and continue to match the classification information of the key frame images of a plurality of consecutive frames with the predetermined action sequence associated with the key frame image of the i-th frame until the user action recognition result is obtained; When the classification information of the key frame image of the (i + 1)-th frame does not match the predetermined action sequence associated with the key frame image of the i-th frame, discard the predetermined action sequence that does not match; The user action recognition method according to claim 7, characterized by including the above.
9. Before the step of obtaining a predetermined action sequence associated with the classification information of the traversed key frame image of the i-th frame, Match the classification information of the key frame image of the i-th frame with the first classification information in each predetermined action sequence, and use the matched predetermined action sequence as the predetermined action sequence associated with the classification information of the key frame image of the i-th frame; The user action recognition method according to claim 8, characterized by including the above.
10. Before traversing the classification information of the key frame image of each frame in the screen recording data, further, Screen the key frame images based on the classification information, and use the screened key frame images as the key frame images to be traversed; The user action recognition method according to claim 1, characterized by including the above.
11. The step of screening the key frame images based on the classification information is Determining whether there are a plurality of consecutive keyframe images having the same classification information according to the time series order of the keyframe images; When there are a plurality of consecutive keyframe images having the same classification information, retaining one keyframe image among the plurality of consecutive keyframe images having the same classification information; The user behavior recognition method according to claim 10, characterized by including the above.
12. The user behavior recognition result includes the duration of the user behavior and / or auxiliary information of the user behavior. The duration of the user behavior is obtained by taking the total duration of the plurality of consecutive keyframe images as the duration of the user behavior. The auxiliary information of the user behavior is obtained by identifying candidate keyframe images from the plurality of keyframe images according to the classification information of the keyframe images of each frame based on the operation actions that are the auxiliary information acquisition targets in the predetermined action sequence, and obtaining the auxiliary information based on the feature information of the candidate keyframe images. The user behavior recognition method according to claim 7, characterized by the above.
13. Obtaining the predetermined action sequence by obtaining the application identifier corresponding to the screen recording data from the log information of the screen recording data and obtaining the predetermined action sequence associated with the application identifier. The user behavior recognition method according to claim 7, characterized by the above.
14. A user behavior recognition device based on screen recording data, A keyframe image extraction unit that extracts keyframe images from a plurality of frame images in the screen recording data to obtain a plurality of keyframe images; A feature information extraction unit that performs data analysis on the keyframe images of each frame to extract the feature information of the keyframe images of each frame; A classification unit that performs screen classification on the keyframe images according to the feature information of the keyframe images of each frame to obtain the classification information of the keyframe images of each frame, where the classification information is a classification unit that characterizes the operation actions performed by the user in the keyframe images. A user behavior recognition unit that traverses the classification information of the key frame images of each frame in the screen recording data and obtains a user behavior recognition result based on the relationship between the classification information of the key frame images of a plurality of consecutive frames. A user behavior recognition device, characterized by including the above.
15. A computer-readable storage medium storing a computer program, When the computer program is executed by a processor, it executes the steps of the user behavior recognition method using the screen recording data according to any one of Claims 1 to 13. A computer-readable storage medium, characterized by the above.
16. A user behavior recognition device including a processor and a memory storing a computer program executable on the processor, When the processor executes the computer program, it executes the steps of the user behavior recognition method using the screen recording data according to any one of Claims 1 to 13. A user behavior recognition device, characterized by the above.
Citation Information
Patent Citations
An analysis system and method for user behavior data based on intelligent image recognition
CN110851148B
Behavioral event measurement system and related methods
JP2017510910A
Image processing device, image processing method and image processing program
JP2020027448A
terminal monitor and program for terminal monitor
JP4069149B1
User behavior tracing method, apparatus and system used in touch screen terminals
US20120281080A1