A resource dispensing determination method and system based on personnel behavior
By acquiring personnel video and voice information, combined with job and work information, and using motion and voice recognition technology to automatically determine resource quantities, the problem of low efficiency and error-proneness in traditional resource allocation is solved, achieving efficient and accurate resource allocation.
Patent Information
- Application Number
- CN202411956138.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-28
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2044-12-28
AI Technical Summary
Traditional methods for determining resource allocation are inefficient and prone to errors, and cannot effectively combine multi-dimensional information for accurate judgment.
By acquiring personnel video and voice information, combined with job information, work content and results information, action and voice recognition technology is used to determine action and voice matching information, and resource quantity is automatically determined based on multi-dimensional information correction coefficients.
It has enabled the automation and multi-dimensional accurate determination of resource quantities, improving the efficiency and accuracy of resource allocation.
Smart Images

Figure CN119918860B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a resource allocation determination method and system based on personnel behavior. BACKGROUND
[0002] In the management and operation of modern enterprises and organizations, the reasonable allocation of resources is of great significance to motivate employees and ensure work efficiency. Traditional resource allocation determination is often based on a single factor, such as only based on work results or fixed post standards.
[0003] In related technologies, resource allocation is usually manually calculated by financial personnel of enterprises and organizations, but this way is not only low in efficiency, but also has a high probability of error.
[0004] Therefore, how to improve the efficiency and accuracy of resource allocation is a research hotspot. SUMMARY
[0005] The embodiments of the present application provide a resource allocation determination method and system based on personnel behavior, which can improve the efficiency and accuracy of resource allocation. The technical solutions are as follows:
[0006] On the one hand, a resource allocation determination method based on personnel behavior is provided, which comprises:
[0007] Obtaining personnel video, personnel information, post information, post resource allocation information and work information of a target work of a target personnel on a target date, the post information is used to describe the post responsibility of the post where the target personnel is located, the post resource allocation information is used to describe the resource allocation standard of the target personnel, and the work information includes work content information and work result information;
[0008] Performing action recognition and voice recognition on the personnel video to determine action information and voice information of the target personnel when working;
[0009] Based on the action information, voice information and work content information of the target personnel when working, determining action matching information and voice matching information of the target personnel, the action matching information is used to represent the matching degree of the action information and the work content information, and the voice matching information is used to represent the matching degree of the voice information and the work content information;
[0010] Based on the post information, the personnel information, the post resource allocation information, the work result information, the action matching information and the voice matching information, determining a target resource amount to be allocated to the target personnel on the target date.
[0011] In one aspect, a personnel behavior-based resource allocation decision system is provided, the system comprising:
[0012] an acquisition module configured to acquire a personnel video, personnel information, post information, post resource allocation information, and work information of a target work of a target personnel on a target date, the post information being configured to describe post responsibilities of a post where the target personnel is located, the post resource allocation information being configured to describe resource allocation standards of the target personnel, the work information comprising work content information and work achievement information;
[0013] an identification module configured to perform action identification and speech identification on the personnel video to determine action information and speech information of the target personnel when working;
[0014] a matching information determination module configured to determine action matching information and speech matching information of the target personnel based on the action information and the speech information of the target personnel when working and the work content information, the action matching information being configured to represent a matching degree between the action information and the work content information, and the speech matching information being configured to represent a matching degree between the speech information and the work content information;
[0015] a resource amount determination module configured to determine a target resource amount to be allocated to the target personnel on the target date based on the post information, the personnel information, the post resource allocation information, the work achievement information, the action matching information, and the speech matching information.
[0016] In one possible implementation, the action information comprises a plurality of actions of the target personnel, the speech information comprises a plurality of speech contents, and the work content information comprises a work type, a work project, and a work target. The matching information determination module is configured to determine first action matching information and second action matching information of the target personnel based on the plurality of actions, the work type, and the work target, the first action matching information being configured to represent a matching degree between the plurality of actions and the work type, and the second action matching information being configured to represent a matching degree between the plurality of actions and the work target; determine action matching information of the target personnel based on the first action matching information and the second action matching information; determine first speech matching information and second speech matching information of the target personnel based on the plurality of speech contents, the work type, and the work project, the first speech matching information being configured to represent a matching degree between the plurality of speech contents and the work type, and the second speech matching information being configured to represent a matching degree between the plurality of speech contents and the work project; and determine speech matching information of the target personnel based on the first speech matching information and the second speech matching information.
[0017] In a possible implementation, the matching information determination module is configured to cluster the actions that are adjacent in time sequence in the plurality of actions to obtain a plurality of action sets; perform feature extraction on the plurality of action sets to obtain action set features of each of the action sets; determine first action matching information of the target personnel based on the action set features of each of the action sets and the work type; and determine second action matching information of the target personnel based on the action set features of each of the action sets and the work target.
[0018] The matching information determination module is configured to cluster the speech contents that are adjacent in time sequence in the plurality of speech contents to obtain a plurality of speech content segments; perform feature extraction on the plurality of speech content segments to obtain speech content segment features of each of the speech content segments; determine first speech content matching information of the target personnel based on the speech content segment features of each of the speech content segments and the work type; and determine second speech content matching information of the target personnel based on the speech content segment features of each of the speech content segments and the work item.
[0019] In a possible implementation, the matching information determination module is configured to encode the first action matching information and the second action matching information based on an attention mechanism to obtain first action matching features of the first action matching information and second action matching features of the second action matching information; fuse the first action matching features and the second action matching features to obtain first fused matching features; and perform multi-round iterative decoding on the first fused matching features based on the attention mechanism to obtain the action matching information of the target personnel.
[0020] The matching information determination module is configured to encode the first speech matching information and the second speech matching information based on an attention mechanism to obtain first speech matching features of the first speech matching information and second speech matching features of the second speech matching information; fuse the first speech matching features and the second speech matching features to obtain second fused matching features; and perform multi-round iterative decoding on the second fused matching features based on the attention mechanism to obtain the speech matching information of the target personnel.
[0021] In a possible implementation, the resource quantity determining module is configured to determine an initial resource quantity based on the post information, the post resource distribution information, and the work result information; determine a first resource quantity correction coefficient based on the post resource distribution information, the action matching information, and the personnel information; determine a second resource quantity correction coefficient based on the post resource distribution information, the voice matching information, and the personnel information; determine a third resource quantity correction coefficient based on the action matching information and the voice matching information; and correct the initial resource quantity by using the first resource quantity correction coefficient, the second resource quantity correction coefficient, and the third resource quantity correction coefficient to obtain a target resource quantity to be distributed to the target personnel on the target date.
[0022] In a possible implementation, the resource quantity determining module is configured to determine result resource description information of the post based on the post information and the post resource distribution information, where the result resource description information is used to indicate resource distribution standards corresponding to different results of completing the target work by the post; and determine the initial resource quantity based on the result resource description information and the work result information.
[0023] In a possible implementation, the personnel information includes physiological parameters and personnel attributes, and the resource quantity determining module is configured to determine a first reference correction coefficient based on the post resource distribution information and the action matching information; determine a second reference correction coefficient based on the action matching information and the personnel information; and fuse the first reference correction coefficient and the second reference correction coefficient to obtain the first resource quantity correction coefficient.
[0024] In a possible implementation, the resource quantity determining module is configured to determine a third reference correction coefficient based on the post resource distribution information and the voice matching information; determine a fourth reference correction coefficient based on the voice matching information and the personnel information; and fuse the third reference correction coefficient and the fourth reference correction coefficient to obtain the second resource quantity correction coefficient.
[0025] In a possible implementation, the resource quantity determining module is configured to perform feature extraction on the action matching information and the voice matching information to obtain action matching features of the action matching information and voice matching features of the voice matching information; fuse the action matching features and the voice matching features to obtain target matching features; and perform full connection and normalization on the target matching features to obtain the third resource quantity correction coefficient.
[0026] In an aspect, a computer device is provided, which comprises one or more processors and one or more memories, and at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the personnel behavior-based resource allocation determination method.
[0027] In an aspect, a computer readable storage medium is provided, which stores at least one computer program, and the computer program is loaded and executed by a processor to implement the personnel behavior-based resource allocation determination method.
[0028] In an aspect, a computer program product or computer program is provided, which comprises program code stored in a computer readable storage medium, and a processor of a computer device reads the program code from the computer readable storage medium, and the processor executes the program code to enable the computer device to perform the personnel behavior-based resource allocation determination method.
[0029] Through the technical scheme provided by the embodiments of the present application, the personnel video, personnel information, post information, post resource allocation information, and work information of the target work of the target personnel on the target date are obtained. The personnel video is subjected to action recognition and voice recognition to determine the action information and voice information of the target personnel when working. Based on the action information and voice information of the target personnel when working and the work content information, the action matching information and voice matching information of the target personnel are determined. Based on the post information, the personnel information, the post resource allocation information, the work achievement information, the action matching information, and the voice matching information, the target resource amount to be allocated to the target personnel on the target date is determined, the automatic determination of the target resource amount is realized, and the accuracy of the determined target resource amount is improved by combining multiple dimensions of information in the determination process. BRIEF DESCRIPTION OF DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0031] Figure 1 is a schematic diagram of an implementation environment of a personnel behavior-based resource allocation determination method provided by an embodiment of the present application;
[0032] Figure 2 is a flowchart of a personnel behavior-based resource allocation determination method provided by an embodiment of the present application;
[0033] Figure 3 is another personnel behavior-based resource allocation determination method flowchart provided by an embodiment of the present application;
[0034] Figure 4 is a structural schematic diagram of a personnel behavior-based resource allocation determination system provided by an embodiment of the present application;
[0035] Figure 5 is a structural schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION
[0036] To make the purposes, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0037] In the present application, the terms "first", "second", and the like are used to distinguish the same or similar items with basically the same function and action, and it should be understood that there is no logical or time sequence dependency between "first", "second", and "nth", and the quantity and execution order are not limited.
[0038] Artificial intelligence (AI) is the use of digital computers or machine controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain better results. Theory, method, technology and application system.
[0039] Machine learning (ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other subjects. It is a branch of computer science that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge sub-models to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence.
[0040] Action recognition: Action recognition is an important branch of computer vision, which mainly studies how to classify human actions in video or image sequences into specific action categories by analyzing them. This technology has wide applications in video surveillance, human-computer interaction, sports analysis, medical rehabilitation and other fields.
[0041] Semantic recognition: Speech recognition, also known as automatic speech recognition (ASR), aims to convert the lexical content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. This technology has wide applications in smart homes, robots, consumer electronics, medical treatment, etc.
[0042] Normalization: mapping a number series with different value ranges to the interval (0, 1) for data processing. In some cases, the normalized numerical value can be directly implemented as a probability.
[0043] Embedded coding: Embedded coding represents a corresponding relationship in mathematics, that is, mapping data on X space to Y space through a function F, where the function F is a single function, and the mapping result is structure preserving. Single function means that the mapped data is uniquely corresponding to the pre-mapped data, and structure preserving means that the size relationship of the pre-mapped data is the same as that of the post-mapped data, for example, there are data X1 and X2 before mapping, and Y1 corresponding to X1 and Y2 corresponding to X2 after mapping. If the data X1 > X2 before mapping, then the data Y1 > Y2 after mapping accordingly. For words, it means mapping words to another space for subsequent machine learning and processing.
[0044] Attention weight: can represent the importance of certain data in the training or prediction process. Importance represents the size of the influence of input data on output data. The data with high importance has a higher value of corresponding attention weight, and the data with low importance has a lower value of corresponding attention weight. In different scenarios, the importance of data is not the same, and the process of training attention weight of the model is to determine the importance of data.
[0045] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0046] In the related art, the amount of resources allocated to personnel is usually manually calculated by the financial personnel of the enterprise, and the efficiency of manual calculation is low, and errors are inevitable in manual calculation, which may cause the resource amount of the resources allocated to the personnel to have a probability of error. By using the technical solutions provided in the embodiments of the present application, the efficiency of resource amount determination can be improved while improving the efficiency of resource amount determination.
[0047] Figure 1is a schematic diagram of an implementation environment of a resource allocation determination method based on personnel behavior provided by an embodiment of the present application, referring to Figure 1 The implementation environment can include a terminal 110 and a server 140.
[0048] The terminal 110 is connected to the server 140 through a wireless network or a wired network. Optionally, the terminal 110 is a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The terminal 110 is installed and runs an application program supporting the resource allocation determination based on personnel behavior. The terminal 110 can acquire relevant information for collecting a target resource amount, and send the collected relevant information to the server 140.
[0049] The server 140 is a stand-alone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and big data and artificial intelligence platform. The server 140 can provide background services for the application program running on the terminal 110.
[0050] The application scenario of the technical solution provided by an embodiment of the present application is described below. The technical solution provided by an embodiment of the present application can be applied in a scenario of determining personnel wages, and can also be applied in a scenario of determining the resource amount of virtual resources allocated to personnel. The present application is not limited thereto.
[0051] A resource allocation determination method based on personnel behavior provided by an embodiment of the present application is described below. Figure 2 is a flowchart of a resource allocation determination method based on personnel behavior provided by an embodiment of the present application, referring to Figure 2 Taking the server as an example of the execution subject, the method includes the following steps.
[0052] 201. The server acquires personnel video, personnel information, post information, post resource allocation information, and work information of a target work of a target personnel on a target date when participating in the target work. The post information is used to describe the post responsibilities of the post where the target personnel is located. The post resource allocation information is used to describe the resource allocation standard of the target personnel. The work information includes work content information and work achievement information.
[0053] Wherein, the target personnel refers to an employee individual selected as a resource quantity determination object within an enterprise or organization. The target date is a pre-determined specific date related to work achievement evaluation. The target work refers to a specific work task or project, which is the basis for work achievement evaluation and resource allocation of the target personnel. The personnel video is a video recorded when the target personnel participates in the target work on the target date. The personnel information is used to describe the basic situation of the target personnel. The post information is used to describe the post responsibilities, which covers detailed description of work tasks, regulations of work processes, required skills and qualification requirements, etc. Taking a software development post as an example, the post responsibilities may include code writing using specific programming languages (such as Java, Python, etc.), following specific software development processes (such as agile development process), having database management skills, etc. The post resource allocation information is used for resource allocation standards. In the case of salary allocation, the post resource allocation information includes salary structure (basic salary, performance salary, etc.), bonus calculation method (based on project completion, sales performance, etc.), welfare benefits (insurance, leave, training opportunities, etc.), etc. The work content information includes work type (such as R&D, sales, customer service, etc.), work project (such as R&D project of specific product, sales project of specific region, etc.), and work goal (such as function implementation goal of R&D project, sales amount goal of sales project, etc.). The work achievement information includes actual achievement records of the target personnel in the target work, such as completed code function modules, achieved sales amount, etc.
[0054] 202. The server performs action recognition and speech recognition on the personnel video to determine action information and speech information of the target personnel during work.
[0055] Wherein, the action information includes multiple actions, and the speech information includes multiple speech contents.
[0056] 203. The server determines action matching information and speech matching information of the target personnel based on the action information and the speech information of the target personnel during work and the work content information, the action matching information being used to represent the matching degree between the action information and the work content information, and the speech matching information being used to represent the matching degree between the speech information and the work content information.
[0057] Wherein, the action matching information is used to represent the matching degree between the action information and the work content information, which can also represent the matching degree between the actions performed by the target personnel during participation in the target work and the work content of the target work. The speech matching information is used to represent the matching degree between the speech information and the work content information, which can also represent the matching degree between the speech of the target personnel during participation in the target work and the work content of the target work.
[0058] 204、the server determines a target resource amount to be distributed to the target personnel on the target date based on the post information, the personnel information, the post resource distribution information, the work achievement information, the action matching information, and the voice matching information.
[0059] The target resource amount is a resource to be distributed to the target personnel, and the resource is salary or virtual resource, which is not limited in the embodiments of the present application.
[0060] According to the technical scheme provided by the embodiments of the present application, the personnel video, the personnel information, the post information, the post resource distribution information, and the work information of the target work of the target personnel on the target date are obtained. The action information and the voice information of the target personnel on the work are determined by performing action recognition and voice recognition on the personnel video. The action matching information and the voice matching information of the target personnel are determined based on the action information and the voice information of the target personnel on the work and the work content information. The target resource amount to be distributed to the target personnel on the target date is determined based on the post information, the personnel information, the post resource distribution information, the work achievement information, the action matching information, and the voice matching information, so as to realize automatic determination of the target resource amount, and the accuracy of the determined target resource amount is improved by combining information of multiple dimensions in the determination process.
[0061] The above steps 201-204 are a brief introduction to the resource distribution determination method based on personnel behavior provided by the embodiments of the present application. The resource distribution determination method based on personnel behavior provided by the embodiments of the present application will be described more clearly in combination with some examples, which are described with reference to Figure 3 Taking the server as an execution subject, the method includes the following steps.
[0062] 301、the server obtains personnel video, personnel information, post information, post resource distribution information, and work information of a target work of a target personnel on a target date, the post information is used to describe post responsibilities of a post where the target personnel is located, the post resource distribution information is used to describe resource distribution standards of the target personnel, and the work information includes work content information and work achievement information.
[0063] The target personnel refers to an employee selected as a resource quantity determination object within an enterprise or organization. The target date is a specific date related to work achievement evaluation. The target work refers to a specific work task or project, which is the basis for work achievement evaluation and resource allocation of the target personnel. The personnel video is a video recorded by the target personnel when participating in the target work on the target date. The personnel information describes the basic situation of the target personnel. The post information describes the post responsibilities, including detailed descriptions of work tasks, regulations of work processes, required skills and qualifications, etc. Taking a software development post as an example, the post responsibilities may include code writing using specific programming languages (such as Java, Python, etc.), following specific software development processes (such as agile development processes), and having database management skills. The post resource allocation information is the resource allocation standard. In the case of salary allocation, the post resource allocation information includes salary structure (basic salary, performance salary, etc.), bonus calculation method (based on project completion, sales performance, etc.), welfare benefits (insurance, leave, training opportunities, etc.), etc. The work content information includes work type (such as R&D, sales, customer service, etc.), work project (such as a specific product development project, a specific region sales project, etc.), and work goal (such as a function implementation goal of a development project, a sales amount goal of a sales project, etc.). The work achievement information includes the actual achievement records of the target personnel in the target work, such as completed code function modules, achieved sales amounts, etc.
[0064] In a possible implementation, the server queries the video database using the personnel identifier of the target personnel and the target date to obtain the personnel video of the target personnel, and the video database stores personnel videos of multiple subjects participating in work on different dates. The server queries the enterprise's human resource management system using the personnel identifier of the target personnel to obtain the personnel information and the post information of the target personnel. The server queries the enterprise's resource management system using the personnel identifier of the target personnel to obtain the post resource allocation information of the target personnel. The server queries the project management system using the work identifier of the target work and the personnel identifier of the target personnel to obtain the work information of the target work.
[0065] 302. The server performs action recognition on the personnel video to determine the action information of the target personnel in the work.
[0066] The action information includes multiple actions of the target personnel.
[0067] In a possible implementation, the server inputs the plurality of video frames of the personnel video into an action recognition model, extracts features of the plurality of video frames of the personnel video by the action recognition model, and obtains video frame features of each video frame. The server determines a plurality of actions of the target personnel based on the video frame features of the plurality of video frames by the action recognition model, to obtain the action information.
[0068] The action recognition model is trained by using a plurality of sample personnel videos and action labels corresponding to each sample personnel video, and has the capability of action recognition. In some embodiments, the action recognition model is ResNet-50 or Inception-v3, etc.
[0069] In this implementation, the action recognition model can be used to recognize the action of the personnel video, and the efficiency and accuracy of action recognition are relatively high.
[0070] For example, the server inputs the plurality of video frames of the personnel video into an action recognition model, extracts features of each video frame of the personnel video by the action recognition model, and obtains video frame features of each video frame. The server obtains the action of each video frame by full connection and normalization of the video frame features of each video frame by the action recognition model, to obtain the action information.
[0071] For example, the server frames the personnel video every preset time length, to obtain the plurality of video frames of the personnel video. The server inputs the plurality of video frames of the personnel video into an action recognition model, performs multi-round convolution on each video frame of the personnel video by the action recognition model, and obtains video frame features of each video frame. The server obtains the action of each video frame by full connection and normalization of the video frame features of each video frame by the action recognition model, to obtain the action information.
[0072] The preset time length is set by a technician according to actual conditions, for example, 1 second or 1.5 seconds, etc., which is not limited in the embodiments of the present application.
[0073] 303. The server performs voice recognition on the personnel video, and determines the voice information of the target personnel at work.
[0074] The voice information includes a plurality of voice contents.
[0075] In a possible implementation, the server extracts personnel audio in the personnel video. The server inputs the personnel audio into a voice recognition model, performs voice recognition on the personnel audio by the voice recognition model, obtains a plurality of voice contents of the target personnel at work, and obtains the voice information.
[0076] The voice content is in a text form. The voice recognition model can be a voice recognition model in the related art, which is not limited in the embodiments of the present application.
[0077] 304. The server determines action matching information of the target person based on the action information of the target person when working and the work content information, where the action matching information is used to indicate a matching degree between the action information and the work content information.
[0078] The action matching information is used to indicate the matching degree between the action information and the work content information, and thus can indicate a matching degree between the action performed by the target person when participating in the target work and the work content of the target work.
[0079] In a possible implementation, the action information includes a plurality of actions of the target person, and the work content information includes a work type and a work target. The server determines first action matching information and second action matching information of the target person based on the plurality of actions, the work type and the work target, where the first action matching information is used to indicate a matching degree between the plurality of actions and the work type, and the second action matching information is used to indicate a matching degree between the plurality of actions and the work target. The server determines the action matching information of the target person based on the first action matching information and the second action matching information.
[0080] In this implementation, the work content information and the plurality of actions are processed to obtain the action matching information of the target person, and the accuracy of the action matching information is relatively high.
[0081] In order to more clearly illustrate the above-mentioned embodiments, the following will be divided into several parts to illustrate the above-mentioned embodiments.
[0082] In a first part, the server determines first action matching information and second action matching information of the target person based on the plurality of actions, the work type and the work target.
[0083] In a possible implementation, the server clusters actions that are adjacent in time sequence in the plurality of actions to obtain a plurality of action sets. The server extracts features of the plurality of action sets to obtain action set features of each action set. The server determines the first action matching information of the target person based on the action set features of each action set and the work type. The server determines the second action matching information of the target person based on the action set features of each action set and the work target.
[0084] For example, the server divides a plurality of actions that are adjacent in time and similar to each other into a same action set to obtain the plurality of action sets. The server extracts features of the plurality of action sets by using an action feature extraction model to obtain action set features of the action sets. The server encodes the work type and the work target to obtain a work type feature of the work type and a work target feature of the work target. The server determines feature similarities between the action set features of the action sets and the work type feature to obtain a plurality of first feature similarities, one first feature similarity corresponding to one action set. The server fuses the plurality of first feature similarities by weighting to obtain first action matching information of the target person. The server determines feature similarities between the action set features of the action sets and the work target feature to obtain a plurality of second feature similarities, one second feature similarity corresponding to one action set. The server fuses the plurality of second feature similarities by weighting to obtain second action matching information of the target person.
[0085] The action feature extraction model is a model trained by a deep neural network (such as an autoencoder). The weights of the weighted fusion are associated with action types corresponding to the action sets.
[0086] The second part, the server determines the action matching information of the target person based on the first action matching information and the second action matching information.
[0087] In a possible implementation, the server encodes the first action matching information and the second action matching information based on an attention mechanism to obtain first action matching features of the first action matching information and second action matching features of the second action matching information. The server fuses the first action matching features and the second action matching features to obtain first fused matching features. The server iteratively decodes the first fused matching features based on the attention mechanism for multiple rounds to obtain the action matching information of the target person.
[0088] For example, the number of heads is set to 8 according to the complexity and performance requirements of the model by using the multi-head attention mechanism. The server encodes the first action matching information and the second action matching information as inputs through the multi-head attention mechanism to obtain the first action matching feature of the first action matching information and the second action matching feature of the second action matching information. The server fuses the first action matching feature and the second action matching feature, for example, by using weighted addition, and the determination of the weight can be based on experimental results or business requirements. If it is found through experiments that the first action matching information contributes more to the overall action matching, it can be given a higher weight, such as 0.6, and the weight of the second action matching information is 0.4. After obtaining the first fused matching feature, the server iteratively decodes the first fused matching feature based on the attention mechanism to finally obtain the action matching information of the target person.
[0089] 305、The server determines speech matching information based on the speech information of the target person during work and the work content information, where the speech matching information is used to represent the matching degree between the speech information and the work content information.
[0090] The speech matching information is used to represent the matching degree between the speech information and the work content information, and thus can represent the matching degree between the speech of the target person during the target work and the work content of the target work.
[0091] In a possible implementation, the speech information includes a plurality of speech contents, the work content information includes a work type and a work item, and the server determines first speech matching information and second speech matching information of the target person based on the plurality of speech contents, the work type, and the work item, where the first speech matching information is used to represent the matching degree between the plurality of speech contents and the work type, and the second speech matching information is used to represent the matching degree between the plurality of speech contents and the work item. The server determines the speech matching information of the target person based on the first speech matching information and the second speech matching information.
[0092] In order to more clearly illustrate the above-mentioned embodiments, the following will be divided into several parts to illustrate the above-mentioned embodiments.
[0093] In the first part, the server determines first speech matching information and second speech matching information of the target person based on the plurality of speech contents, the work type, and the work item.
[0094] In a possible implementation, the server clusters the plurality of voice contents that are adjacent in time sequence to obtain a plurality of voice content segments. The server extracts features of the plurality of voice content segments to obtain voice content segment features of each voice content segment. The server determines first voice content matching information of the target personnel based on the voice content segment features of each voice content segment and the work type. The server determines second voice content matching information of the target personnel based on the voice content segment features of each voice content segment and the work project.
[0095] For example, the server divides the plurality of voice contents that are adjacent in time sequence and similar into the same voice content segment to obtain the plurality of voice content segments. The server extracts features of the plurality of voice content segments by using a voice content feature extraction model to obtain voice content segment features of each voice content segment. The server performs embedding coding on the work type and the work project to obtain work type features of the work type and work project features of the work project. The server determines feature similarities between the voice content segment features of each voice content segment and the work type features to obtain a plurality of third feature similarities, one third feature similarity corresponding to one voice content segment. The server performs weighted fusion on the plurality of third feature similarities to obtain the first voice content matching information of the target personnel. The server determines feature similarities between the voice content segment features of each voice content segment and the work project features to obtain a plurality of fourth feature similarities, one fourth feature similarity corresponding to one voice content segment. The server performs weighted fusion on the plurality of fourth feature similarities to obtain the second voice content matching information of the target personnel.
[0096] The voice content feature extraction model is a model trained by a deep neural network (such as an autoencoder). The weights of the weighted fusion are associated with voice content types corresponding to the voice content segments.
[0097] The second part, the server determines the voice matching information of the target personnel based on the first voice matching information and the second voice matching information.
[0098] In a possible implementation, the server encodes the first voice matching information and the second voice matching information based on an attention mechanism to obtain first voice matching features of the first voice matching information and second voice matching features of the second voice matching information. The server fuses the first voice matching features and the second voice matching features to obtain second fused matching features. The server iteratively decodes the second fused matching features based on the attention mechanism for multiple rounds to obtain the voice matching information of the target personnel.
[0099] For example, the server adopts a multi-head attention mechanism, and the number of heads is set to 4. The first voice matching information and the second voice matching information are respectively taken as inputs, and after encoding by the multi-head attention mechanism, the first voice matching feature of the first voice matching information and the second voice matching feature of the second voice matching information are obtained. The server fuses the first voice matching feature and the second voice matching feature, for example, in a weighted multiplication manner, and the weight is determined according to experiments or business requirements. If it is found through analysis that the first voice matching information has a relatively small influence on the final voice matching, a lower weight, such as 0.3, can be given to it, and the weight of the second voice matching information is 0.7. After obtaining the second fused matching feature, the second fused matching information is iteratively decoded based on the attention mechanism, and finally the voice matching information of the target person is obtained.
[0100] 306. The server determines a target resource amount to be distributed to the target person on the target date based on the post information, the personnel information, the post resource distribution information, the work achievement information, the action matching information, and the voice matching information.
[0101] The target resource amount is a resource distributed to the target person, and the resource can be salary or virtual resource, which is not limited by the embodiments of the present application. In the case of salary, the target resource amount is the numerical value of the salary, and in the case of virtual resource (such as ticket or shopping card), the target resource amount is the resource amount of the virtual resource. In the above step 306, the target resource amount determined is the resource amount of the target date, that is, the resource amount determination with a granularity of day is realized, and in combination with a longer time range, the target resource amount corresponding to the longer time range can be obtained.
[0102] In a possible implementation, the server determines an initial resource amount based on the post information, the post resource distribution information, and the work achievement information. The server determines a first resource amount correction coefficient based on the post resource distribution information, the action matching information, and the personnel information. The server determines a second resource amount correction coefficient based on the post resource distribution information, the voice matching information, and the personnel information. The server determines a third resource amount correction coefficient based on the action matching information and the voice matching information. The server corrects the initial resource amount by using the first resource amount correction coefficient, the second resource amount correction coefficient, and the third resource amount correction coefficient to obtain the target resource amount to be distributed to the target person on the target date.
[0103] The initial resource amount is a reference resource amount, and the first correction coefficient, the second correction coefficient, and the third correction coefficient are all correction coefficients for correcting the initial resource amount, thereby improving the accuracy of the target resource amount.
[0104] To make the above-mentioned embodiments more clearly, the following is divided into several parts to explain the above-mentioned embodiments.
[0105] The first part, the server determines the initial resource amount based on the post information, the post resource allocation information and the work achievement information.
[0106] In a possible implementation, the server determines the achievement resource description information of the post based on the post information and the post resource allocation information, the achievement resource description information being used to represent resource allocation standards corresponding to different achievements of the post in completing the target work. The server determines the initial resource amount based on the achievement resource description information and the work achievement information.
[0107] For example, the server performs semantic feature extraction on the post information and the post resource allocation information to obtain first semantic features of the post information and second semantic features of the post resource allocation information. The server fuses the first semantic features and the second semantic features to obtain first semantic fusion features. The server performs multi-round iterative decoding on the first semantic fusion features based on an attention mechanism to obtain achievement resource description information. The server queries the work achievement information in the achievement resource description information to obtain the initial resource amount corresponding to the work achievement information.
[0108] The second part, the server determines the first resource amount correction coefficient based on the post resource allocation information, the action matching information and the personnel information.
[0109] In a possible implementation, the personnel information includes physiological parameters and personnel attributes, the server determines a first reference correction coefficient based on the post resource allocation information and the action matching information. The server determines a second reference correction coefficient based on the action matching information and the personnel information. The server fuses the first reference correction coefficient and the second reference correction coefficient to obtain the first resource amount correction coefficient.
[0110] The physiological parameters include heart rate, blood pressure and body temperature and the like when the target personnel participates in the target work.
[0111] For example, the server extracts semantic features of the post resource allocation information to obtain third semantic features of the post resource allocation information. The server fuses the third semantic features and the action matching information, and then performs full connection and normalization to obtain the first reference correction coefficient. The server extracts features of the physiological parameters and the personnel attributes to obtain physiological parameter features of the physiological parameters and personnel attribute features of the personnel attributes. The server fuses the physiological parameter features, the personnel attribute features, and the action matching information, and then performs full connection and normalization to obtain the second reference correction coefficient. The server fuses the first reference correction coefficient and the second reference correction coefficient by weighting to obtain the first resource quantity correction coefficient.
[0112] The weights of the weighting fusion are set by technicians according to actual conditions, and embodiments of the application are not limited in this regard.
[0113] In a possible implementation, the server determines a third reference correction coefficient based on the post resource allocation information and the voice matching information. The server determines a fourth reference correction coefficient based on the voice matching information and the personnel information. The server fuses the third reference correction coefficient and the fourth reference correction coefficient to obtain the second resource quantity correction coefficient.
[0114] In a possible implementation, the server determines a third reference correction coefficient based on the post resource allocation information and the voice matching information. The server determines a fourth reference correction coefficient based on the voice matching information and the personnel information. The server fuses the third reference correction coefficient and the fourth reference correction coefficient to obtain the second resource quantity correction coefficient.
[0115] For example, the server extracts semantic features of the post resource allocation information to obtain third semantic features of the post resource allocation information. The server fuses the third semantic features and the voice matching information, and then performs full connection and normalization to obtain the third reference correction coefficient. The server extracts features of the physiological parameters and the personnel attributes to obtain physiological parameter features of the physiological parameters and personnel attribute features of the personnel attributes. The server fuses the physiological parameter features, the personnel attribute features, and the voice matching information, and then performs full connection and normalization to obtain the fourth reference correction coefficient. The server fuses the third reference correction coefficient and the fourth reference correction coefficient by weighting to obtain the second resource quantity correction coefficient.
[0116] The weights of the weighting fusion are set by technicians according to actual conditions, and embodiments of the application are not limited in this regard.
[0117] Alternatively, the second resource quantity correction coefficient can also be determined in the following manner.
[0118] In a possible implementation, the server determines a third reference correction coefficient based on the post resource allocation information and the voice matching information. The server determines a fourth reference correction coefficient based on the voice matching information and the personnel information. The server fuses the third reference correction coefficient and the fourth reference correction coefficient to obtain the second resource quantity correction coefficient.
[0119] In a possible implementation, the server performs feature extraction on the action matching information and the speech matching information to obtain action matching features of the action matching information and speech matching features of the speech matching information. The server fuses the action matching features and the speech matching features to obtain target matching features. The server performs full connection and normalization on the target matching features to obtain the third resource quantity correction coefficient.
[0120] In a possible implementation, the server multiplies the first resource quantity correction coefficient, the second resource quantity correction coefficient, and the third resource quantity correction coefficient by the initial resource quantity to obtain the target resource quantity.
[0121] In a possible implementation, the server multiplies the first resource quantity correction coefficient, the second resource quantity correction coefficient, and the third resource quantity correction coefficient by the initial resource quantity to obtain the target resource quantity.
[0122] All the optional technical solutions described above can be combined to form optional embodiments of the present application, which will not be described again.
[0123] Through the technical solutions provided in the embodiments of the present application, personnel video, personnel information, post information, post resource distribution information, and work information of a target work of a target person on a target date are obtained. Action recognition and speech recognition are performed on the personnel video to determine action information and speech information of the target person when working. Based on the action information, the speech information, and the work content information of the target person when working, action matching information and speech matching information of the target person are determined. Based on the post information, the personnel information, the post resource distribution information, the work achievement information, the action matching information, and the speech matching information, a target resource quantity to be distributed to the target person on the target date is determined, which realizes automatic determination of the target resource quantity and improves the accuracy of the determined target resource quantity by combining information in multiple dimensions in the determination process.
[0124] Figure 4 is a structural schematic diagram of a resource distribution determination system based on personnel behavior provided by an embodiment of the present application, referring to Figure 4 The system includes an acquisition module 401, an identification module 402, a matching information determination module 403, and a resource quantity determination module 404.
[0125] The acquisition module 401 is configured to acquire personnel video, personnel information, post information, post resource allocation information, and work information of a target work of a target person on a target date, the post information is used to describe post responsibilities of a post where the target person is located, the post resource allocation information is used to describe resource allocation standards of the target person, and the work information includes work content information and work achievement information.
[0126] The recognition module 402 is configured to perform action recognition and voice recognition on the personnel video, and determine action information and voice information of the target person when working.
[0127] The matching information determination module 403 is configured to determine action matching information and voice matching information of the target person based on the action information, the voice information, and the work content information of the target person when working, the action matching information is used to represent a matching degree between the action information and the work content information, and the voice matching information is used to represent a matching degree between the voice information and the work content information.
[0128] The resource quantity determination module 404 is configured to determine a target resource quantity allocated to the target person on the target date based on the post information, the personnel information, the post resource allocation information, the work achievement information, the action matching information, and the voice matching information.
[0129] In a possible implementation, the action information includes a plurality of actions of the target person, the voice information includes a plurality of voice contents, the work content information includes a work type, a work item, and a work target, and the matching information determination module 403 is configured to determine first action matching information and second action matching information of the target person based on the plurality of actions, the work type, and the work target, the first action matching information is used to represent a matching degree between the plurality of actions and the work type, and the second action matching information is used to represent a matching degree between the plurality of actions and the work target. The action matching information of the target person is determined based on the first action matching information and the second action matching information. The first voice matching information and the second voice matching information of the target person are determined based on the plurality of voice contents, the work type, and the work item, the first voice matching information is used to represent a matching degree between the plurality of voice contents and the work type, and the second voice matching information is used to represent a matching degree between the plurality of voice contents and the work item. The voice matching information of the target person is determined based on the first voice matching information and the second voice matching information.
[0130] In a possible implementation, the matching information determination module 403 is configured to cluster the actions that are adjacent in time sequence in the plurality of actions to obtain a plurality of action sets. Feature extraction is performed on the plurality of action sets to obtain action set features of each of the action sets. The first action matching information of the target person is determined based on the action set features of each of the action sets and the work type. The second action matching information of the target person is determined based on the action set features of each of the action sets and the work target.
[0131] The matching information determination module 403 is configured to cluster the speech contents that are adjacent in time sequence in the plurality of speech contents to obtain a plurality of speech content segments. Feature extraction is performed on the plurality of speech content segments to obtain speech content segment features of each of the speech content segments. The first speech content matching information of the target person is determined based on the speech content segment features of each of the speech content segments and the work type. The second speech content matching information of the target person is determined based on the speech content segment features of each of the speech content segments and the work item.
[0132] In a possible implementation, the matching information determination module 403 is configured to encode the first action matching information and the second action matching information based on an attention mechanism to obtain first action matching features of the first action matching information and second action matching features of the second action matching information. The first action matching features and the second action matching features are fused to obtain first fused matching features. The first fused matching features are iteratively decoded in multiple rounds based on the attention mechanism to obtain the action matching information of the target person.
[0133] The matching information determination module 403 is configured to encode the first speech matching information and the second speech matching information based on an attention mechanism to obtain first speech matching features of the first speech matching information and second speech matching features of the second speech matching information. The first speech matching features and the second speech matching features are fused to obtain second fused matching features. The second fused matching features are iteratively decoded in multiple rounds based on the attention mechanism to obtain the speech matching information of the target person.
[0134] In a possible implementation, the resource quantity determination module 404 is configured to determine an initial resource quantity based on the post information, the post resource allocation information, and the work result information. A first resource quantity correction coefficient is determined based on the post resource allocation information, the action matching information, and the personnel information. A second resource quantity correction coefficient is determined based on the post resource allocation information, the voice matching information, and the personnel information. A third resource quantity correction coefficient is determined based on the action matching information and the voice matching information. The initial resource quantity is corrected by using the first resource quantity correction coefficient, the second resource quantity correction coefficient, and the third resource quantity correction coefficient, to obtain a target resource quantity allocated to the target personnel on the target date.
[0135] In a possible implementation, the resource quantity determination module 404 is configured to determine, based on the post information and the post resource allocation information, result resource description information of the post, where the result resource description information is used to indicate resource allocation standards corresponding to different results of completing the target work by the post. The initial resource quantity is determined based on the result resource description information and the work result information.
[0136] In a possible implementation, the personnel information includes physiological parameters and personnel attributes, and the resource quantity determination module 404 is configured to determine a first reference correction coefficient based on the post resource allocation information and the action matching information. A second reference correction coefficient is determined based on the action matching information and the personnel information. The first reference correction coefficient and the second reference correction coefficient are fused to obtain the first resource quantity correction coefficient.
[0137] In a possible implementation, the resource quantity determination module 404 is configured to determine a third reference correction coefficient based on the post resource allocation information and the voice matching information. A fourth reference correction coefficient is determined based on the voice matching information and the personnel information. The third reference correction coefficient and the fourth reference correction coefficient are fused to obtain the second resource quantity correction coefficient.
[0138] In a possible implementation, the resource quantity determination module 404 is configured to perform feature extraction on the action matching information and the voice matching information, to obtain action matching features of the action matching information and voice matching features of the voice matching information. The action matching features and the voice matching features are fused to obtain target matching features. The target matching features are fully connected and normalized to obtain the third resource quantity correction coefficient.
[0139] It should be noted that the above embodiment provides a resource allocation determination system based on personnel behavior, when determining the resource allocation, only the division of the above functional modules is exemplified, in actual application, the above functional distribution can be completed by different functional modules according to the needs, that is, the internal structure of the server is divided into different functional modules to complete all or part of the functions described above. In addition, the resource allocation determination system based on personnel behavior and the resource allocation determination method based on personnel behavior provided by the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0140] Through the technical solutions provided by the embodiments of the present application, the personnel video, personnel information, post information, post resource allocation information and work information of the target work of the target personnel on the target date are obtained. The action recognition and voice recognition are performed on the personnel video to determine the action information and voice information of the target personnel when working. Based on the action information, voice information and work content information of the target personnel when working, the action matching information and voice matching information of the target personnel are determined. Based on the post information, the personnel information, the post resource allocation information, the work result information, the action matching information and the voice matching information, the target resource amount allocated to the target personnel on the target date is determined, the automatic determination of the target resource amount is realized, and the accuracy of the determined target resource amount is improved by combining multiple dimensions of information in the determination process.
[0141] Figure 5 The server 500 can have a large difference due to different configurations or performances, and can include one or more processors (Central Processing Units, CPUs) 501 and one or more memories 502, wherein at least one computer program is stored in the one or more memories 502, and the at least one computer program is loaded and executed by the one or more processors 501 to realize the method provided by each method embodiment. Of course, the server 500 can also have a wired or wireless network interface, a keyboard, an input and output interface and other components for realizing the functions of the device, so as to perform input and output. The server 500 can also include other components for realizing the functions of the device, which will not be repeated here.
[0142] In the example embodiment, a computer readable storage medium, such as a memory including a computer program executable by a processor to implement one of the above-mentioned personnel behavior-based resource allocation determination methods, is also provided. For example, the computer readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0143] In the example embodiment, a computer program product or computer program including program code stored in a computer readable storage medium is also provided, and the processor of a computer device reads the program code from the computer readable storage medium, and the processor executes the program code to cause the computer device to implement one of the above-mentioned personnel behavior-based resource allocation determination methods.
[0144] In some embodiments, the computer program related to the embodiments of the present application can be deployed to execute on one computer device, or on multiple computer devices located in one place, or on multiple computer devices distributed in multiple places and interconnected through a communication network, which can constitute a blockchain system.
[0145] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing relevant hardware, and the program can be stored in a computer readable storage medium, such as a Read-Only Memory, a magnetic disk or an optical disk, etc.
[0146] The above is only an optional embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A resource allocation determination method based on personnel behavior, characterized in that, The method comprises: acquiring personnel video, personnel information, post information, post resource allocation information and work information of a target work of a target person on a target date, the post information being used to describe post responsibilities of a post where the target person is located, the post resource allocation information being used to describe resource allocation standards of the target person, the work information comprising work content information and work result information, and the personnel information comprising physiological parameters and personnel attributes; performing action recognition and voice recognition on the personnel video to determine action information and voice information of the target person when working; based on the action information, the voice information and the work content information of the target person when working, determining action matching information and voice matching information of the target person, the action matching information being used to represent a matching degree of the action information and the work content information, the voice matching information being used to represent a matching degree of the voice information and the work content information, the action information comprising multiple actions of the target person, and the voice information comprising multiple voice contents; based on the post information, the post resource allocation information and the work result information, determining an initial resource amount; performing semantic feature extraction on the post resource allocation information to obtain third semantic features of the post resource allocation information; performing full connection and normalization on the third semantic features and the action matching information after fusion to obtain a first reference correction coefficient; performing feature extraction on the physiological parameters and the personnel attributes to obtain physiological parameter features of the physiological parameters and personnel attribute features of the personnel attributes; performing full connection and normalization on the physiological parameter features, the personnel attribute features and the action matching information after fusion to obtain a second reference correction coefficient; and performing weighted fusion on the first reference correction coefficient and the second reference correction coefficient to obtain a first resource amount correction coefficient; performing full connection and normalization on the third semantic features and the voice matching information after fusion to obtain a third reference correction coefficient; performing feature extraction on the physiological parameters and the personnel attributes to obtain physiological parameter features of the physiological parameters and personnel attribute features of the personnel attributes; performing full connection and normalization on the physiological parameter features, the personnel attribute features and the voice matching information after fusion to obtain a fourth reference correction coefficient; and performing weighted fusion on the third reference correction coefficient and the fourth reference correction coefficient to obtain a second resource amount correction coefficient; performing feature extraction on the action matching information and the voice matching information to obtain action matching features of the action matching information and voice matching features of the voice matching information; performing fusion on the action matching features and the voice matching features to obtain target matching features; and performing full connection and normalization on the target matching features to obtain a third resource amount correction coefficient; multiplying the first resource amount correction coefficient, the second resource amount correction coefficient and the third resource amount correction coefficient by the initial resource amount to obtain a target resource amount.
2. The method of claim 1, wherein, The work content information includes work type, work project, and work objective. The step of determining the target person's action matching information and voice matching information based on the target person's action information, voice information, and the work content information includes: Based on the multiple actions, the job type, and the job objective, first action matching information and second action matching information of the target personnel are determined. The first action matching information is used to indicate the degree of matching between the multiple actions and the job type, and the second action matching information is used to indicate the degree of matching between the multiple actions and the job objective. Based on the first action matching information and the second action matching information, determine the action matching information of the target person; Based on the multiple voice contents, the job type, and the job item, first voice matching information and second voice matching information of the target person are determined. The first voice matching information is used to indicate the degree of matching between the multiple voice contents and the job type, and the second voice matching information is used to indicate the degree of matching between the multiple voice contents and the job item. Based on the first voice matching information and the second voice matching information, the voice matching information of the target person is determined.
3. The method of claim 2, wherein, The step of determining the first action matching information and the second action matching information of the target personnel based on the multiple actions, the work type, and the work objective includes: Clustering sequentially adjacent actions among the multiple actions yields multiple action sets; feature extraction is performed on the multiple action sets to obtain action set features for each action set; based on the action set features of each action set and the job type, first action matching information for the target personnel is determined; based on the action set features of each action set and the job target, second action matching information for the target personnel is determined. The step of determining the first voice matching information and the second voice matching information of the target person based on the multiple voice contents, the work type, and the work project includes: Clustering of temporally adjacent speech content from the plurality of speech content segments yields a plurality of speech content segments; feature extraction is performed on the plurality of speech content segments to obtain speech content segment features for each speech content segment; based on the speech content segment features of each speech content segment and the work type, first speech content matching information for the target personnel is determined; based on the speech content segment features of each speech content segment and the work project, second speech content matching information for the target personnel is determined.
4. The method of claim 2, wherein, The step of determining the action matching information of the target person based on the first action matching information and the second action matching information includes: The first action matching information and the second action matching information are encoded based on the attention mechanism to obtain the first action matching feature of the first action matching information and the second action matching feature of the second action matching information. The first action matching feature and the second action matching feature are fused to obtain first fused matching features; The first fused matching features are iteratively decoded based on an attention mechanism to obtain action matching information of the target personnel; The speech matching information of the target personnel is determined based on the first speech matching information and the second speech matching information, including: The first speech matching information and the second speech matching information are respectively encoded based on an attention mechanism to obtain first speech matching features of the first speech matching information and second speech matching features of the second speech matching information; The first speech matching features and the second speech matching features are fused to obtain second fused matching features; The second fused matching features are iteratively decoded based on an attention mechanism to obtain the speech matching information of the target personnel.
5. The method of claim 1, wherein, The initial resource amount is determined based on the post information, the post resource distribution information, and the work achievement information, including: The achievement resource description information of the post is determined based on the post information and the post resource distribution information, and the achievement resource description information is used to represent resource distribution standards corresponding to different achievements of the target work completed by the post; The initial resource amount is determined based on the achievement resource description information and the work achievement information.
6. A resource distribution determination system based on personnel behavior, characterized in that, an acquisition module is configured to acquire personnel video, personnel information, post information, post resource distribution information, and work information of a target work of a target personnel participating in the target work on a target date, the post information is used to describe post responsibilities of a post where the target personnel is located, the post resource distribution information is used to describe resource distribution standards of the target personnel, the work information includes work content information and work achievement information, and the personnel information includes physiological parameters and personnel attributes; an identification module is configured to perform action recognition and speech recognition on the personnel video to determine action information and speech information of the target personnel during work; a matching information determination module is configured to determine action matching information and speech matching information of the target personnel based on the action information, the speech information of the target personnel during work, and the work content information, the action matching information is used to represent a matching degree between the action information and the work content information, and the speech matching information is used to represent a matching degree between the speech information and the work content information; a resource amount determination module is configured to determine an initial resource amount based on the post information, the post resource distribution information, and the work achievement information; semantic features of the post resource distribution information are extracted to obtain third semantic features of the post resource distribution information; the third semantic features and the action matching information are fused, fully connected, and normalized to obtain a first reference correction coefficient; physiological parameter features of the physiological parameters and personnel attribute features of the personnel attributes are extracted from the physiological parameters and the personnel attributes; The physiological parameter feature, the personnel attribute feature and the action matching information are fused, full connection and normalization are performed, and a second reference correction coefficient is obtained; The first reference correction coefficient and the second reference correction coefficient are weighted and fused to obtain a first resource quantity correction coefficient; The third semantic feature and the voice matching information are fused, full connection and normalization are performed, and a third reference correction coefficient is obtained; the physiological parameter and the personnel attribute are extracted to obtain a physiological parameter feature of the physiological parameter and a personnel attribute feature of the personnel attribute; The physiological parameter feature, the personnel attribute feature and the voice matching information are fused, full connection and normalization are performed, and a fourth reference correction coefficient is obtained; The third reference correction coefficient and the fourth reference correction coefficient are weighted and fused to obtain a second resource quantity correction coefficient; The action matching information and the voice matching information are extracted to obtain an action matching feature of the action matching information and a voice matching feature of the voice matching information; The action matching feature and the voice matching feature are fused to obtain a target matching feature; Full connection and normalization are performed on the target matching feature to obtain a third resource quantity correction coefficient; the first resource quantity correction coefficient, the second resource quantity correction coefficient and the third resource quantity correction coefficient are multiplied by the initial resource quantity to obtain a target resource quantity.
Citation Information
Patent Citations
Person-and-post matching detection method and device, computer equipment and storage medium
CN113435380A
Working state process monitoring method and device and storage medium
CN116629651A
Human resource management method and system
CN116993312A