Employee integral data processing method and system

By using mobile acquisition devices and semantic processing technology, high-quality audio and video capture of offline meetings and accurate extraction and association of employee tasks are achieved, solving the problems of completeness of offline meeting records and task location, and improving employee work efficiency and project execution results.

CN121728217APending Publication Date: 2026-03-24ZHEJIANG XIEDING DIGITAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies cannot capture complete and high-quality audio and video of presentations in offline meetings, nor can they accurately extract and associate personalized task information with employees, making it difficult for employees to quickly locate meeting content relevant to their own tasks.

Method used

Video and audio are captured using mobile acquisition devices. Combined with semantic extraction and video segmentation technologies, the system generates personalized task text and step-by-step image sets for each employee, and updates employee points based on project deadlines.

Benefits of technology

It enabled accurate recording of offline meeting content and rapid location of employee tasks, improving viewing efficiency and work completion motivation, and enhancing project execution results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728217A_ABST
    Figure CN121728217A_ABST
Patent Text Reader

Abstract

The invention provides an employee integral data processing method and system. Belongs to the technical field of data processing. The method comprises the following steps: controlling a video acquisition unit and a voice acquisition unit of a mobile acquisition device to respectively carry out video acquisition and voice acquisition on speech operation of an offline speaker to obtain a recorded video and recorded voice; performing semantic extraction on the recorded voice based on each employee name corresponding to each employee terminal to obtain each task text comprising different step sub-texts; carrying out video splitting on the recorded video based on each task text, and carrying out fragment identification on each obtained task fragment to obtain each step image group comprising different step images; sending the task text and each step image group corresponding to the same employee terminal to the employee terminal, and obtaining step completion data uploaded by each employee terminal based on the project deadline; and performing integral updating on the employee integral of each employee terminal based on each uploading quantity of the data completed in each step. The method at least improves the employee incentive effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to data processing technology, and in particular to an employee credit data processing method and system. BACKGROUND

[0002] In the process of daily operation and project promotion of enterprises, offline meetings are the core scene of information transmission, task allocation and decision synchronization, especially for project meetings involving multi-department cooperation and multi-employee participation. The complete recording of meeting content and the accurate disassembly and tracking of tasks directly affect the project execution efficiency and goal achievement rate. Currently, the recording method of offline meetings mainly relies on manual notes, fixed position camera recording or audio recorder recording. Although some technical solutions can realize the preliminary collection of meeting audio and video, subsequent processing is mostly limited to the storage and playback of audio and video files, and lacks deep analysis of meeting content and task association.

[0003] In the prior art, Chinese patent application No. CN202210876543.2 discloses an "intelligent meeting recording and task allocation system". This scheme collects meeting audio and video through fixed cameras and microphones in the meeting room, converts speech into text records using voice recognition technology, extracts task information from the meeting based on keywords, and then sends the task list to the relevant employee terminal through email. This technology realizes the automation of meeting recording and the preliminary allocation of tasks, and to some extent reduces the workload of manual recording. However, this scheme has obvious defects: it can only extract general task lists, cannot associate personalized task texts for each employee, and cannot correspondingly disassemble tasks and specific steps in the meeting video, making it difficult for employees to quickly locate meeting content related to their own tasks. Based on the above review of the prior art, the core technical problem that needs to be solved in the current industry but has not been broken through by the prior art is: how to realize complete and high-quality audio and video collection of offline meeting content, and at the same time distinguish and extract task information associated with different employees from the meeting speech, and then accurately associate and disassemble these task information with the corresponding content in the meeting video. SUMMARY

[0004] Based on the above problems, the present application is proposed to provide an employee credit data processing method and system that overcomes the above problems or at least partially solves the above problems.

[0005] According to one aspect of the present application, an employee credit data processing method is provided, comprising the following steps: controlling the video collection unit and the voice collection unit of the mobile collection device to respectively collect the video and voice of the offline speaker's speech operation, to obtain the recorded video and recorded voice; Semantic extraction is performed on the recorded audio based on the employee names corresponding to each employee's terminal, resulting in task texts that include sub-texts of different steps. Based on the text of each task, the recorded video is segmented, and the resulting task segments are segmented to obtain a group of images for each step, including images of different steps. Send the task text and image sets of each step corresponding to the same employee's terminal to the employee's terminal, and obtain the step completion data uploaded by each employee's terminal based on the project deadline; The employee points on each employee's end are updated based on the number of data uploaded in each of the aforementioned steps.

[0006] Optionally, in the method according to the present invention, the video acquisition unit and the voice acquisition unit of the mobile acquisition device respectively acquire video and voice data of the offline speaker's presentation, resulting in recorded video and recorded voice data, including: When the location debugging signal sent by the management terminal is received, the video acquisition unit of the mobile acquisition device is controlled to acquire video and obtain debugging video. Perform image comparison based on adjacent positions on each video image frame that makes up the debug video; When it is determined from the comparison results that the two video image frames at the end have the same image content, the mobile acquisition device is controlled to end the video acquisition, and either of the two video image frames is determined as the identification reference frame. When image recognition is performed on the recognition reference frame to determine that the position attribute corresponding to the recognition reference frame is a suitable attribute, the video acquisition unit and voice acquisition unit of the mobile acquisition device are controlled to perform video acquisition and voice acquisition of the offline speaker's speech operation, respectively, to obtain recorded video and recorded voice.

[0007] Optionally, in the method according to the invention, the method further includes: The identification reference frame is binarized, and the resulting binarized image is pixel-recognized to obtain each image pixel that makes up the binarized image. The binarized image includes screen pixels corresponding to the first pixel value and / or noise pixels corresponding to the second pixel value. If each image pixel includes at least one image pixel corresponding to the first pixel value, the at least one image pixel is determined as each edge pixel, and it is determined whether each edge pixel is located on the image contour of the corresponding binarized image; If none of the edge pixels are located within the image contour of the binarized image, the position attribute corresponding to the recognition reference frame is determined to be an appropriate attribute.

[0008] Optionally, in the method according to the invention, the method further includes: If any edge pixel point is located on the image contour of the binary image, determine each image sub-contour constituting the image contour; Determine each edge pixel point located on each image sub-contour and exhibiting continuous adjacency as a same frame contour segment, and obtain each frame contour segment located in the recognition reference frame corresponding to the binary image; Determine the number of line segments corresponding to each frame contour segment, and if the number of line segments is 1, determine the midpoint of the line segment corresponding to the frame contour segment; Generate an offset indication line based on the recognition reference frame, with the midpoint of the line segment as the starting point and pointing to the image center point of the corresponding recognition reference frame, and send the obtained debugging indication image to the management end.

[0009] Optionally, in the method according to the present application, semantic extraction is performed on the recorded voice based on the employee names corresponding to each employee terminal, and each task text including different step subtexts is obtained, including: The preset participant list is called, wherein the preset participant list includes each employee terminal and the employee names corresponding to each employee terminal; Semantic recognition is performed on the recorded voice, and text segmentation is performed on the obtained voice text based on the employee names of the preset participant list, to obtain each task text corresponding to each employee terminal corresponding to each employee name; Step labeling is performed on each task text based on each step keyword corresponding to different preset step templates, to obtain each step subtext corresponding to different task steps.

[0010] Optionally, in the method according to the present application, video splitting is performed on the recorded video based on each task text, and segment recognition is performed on each task segment obtained, to obtain each step image group including different step images, including: The recorded video is video split based on the obtained start time and end time corresponding to each task text, to obtain each task segment corresponding to each employee terminal; Each task segment is video split based on the obtained step time corresponding to each step label, to obtain each step segment corresponding to each step subtext; Each video image frame corresponding to the same step segment and having an adjacent relationship is sequentially subjected to image comparison, and each video image frame having a difference region is determined as a region change image based on the corresponding comparison result; Each region change image corresponding to the same step segment is sequentially filled with a serial number in a preset marking region based on the time sequence, and each step image corresponding to the same step segment is divided into the same step image group.

[0011] Optionally, in the method according to the present application, the video image frames having the adjacent relationship corresponding to the same step segment are sequentially subjected to image comparison, and based on the corresponding comparison result, the region change image of each video image frame with the difference region is determined, including: dividing each video image frame into an upper left sub-region, a lower left sub-region, an upper right sub-region and a lower right sub-region based on the image center point of each video image frame; dividing each video image frame corresponding to the same step segment into the same image determination group, and sequentially performing image comparison between each video image frame located in each image determination group and the video image frame located in front of the video image frame, to obtain each difference region; determining each sub-region corresponding to each difference region as the speech region corresponding to the video image frame; dividing each video image frame corresponding to the same speech region having the adjacent relationship into the same region image group, and determining each video image frame located at the end of each region image group as the region change image.

[0012] Optionally, in the method according to the present application, sequentially performing image comparison between each video image frame located in each image determination group and the video image frame located in front of the video image frame, to obtain each difference region, including: dividing each video image frame located in each image determination group and the video image frame located in front of the video image frame into the same image comparison group; respectively performing binaryzation processing on each video image frame located in each image comparison group, to obtain each first image and each second image corresponding to the same image comparison group, wherein any first image includes each first image pixel point corresponding to a first pixel value, and any second image includes each second image pixel point corresponding to a second pixel value; creating a transparent fusion layer, and sequentially stacking the first image and the second image corresponding to the same image comparison group on the transparent fusion layer; if any second image pixel point is included in the transparent fusion layer, performing pixel connection on the second image pixel points to obtain each difference region.

[0013] Optionally, in the method according to the present application, the method further includes: determining the step image located at the end of each step image group as the comparison image frame, and performing time-based sorting from early to late on each step image group to obtain a comparison sequence; respectively performing image comparison between each step image located in each step image group and the comparison image frame corresponding to the step image group located in front of the step image group based on the comparison sequence, and determining each region part with the image difference as the comparison result as each newly added annotation region corresponding to each step image; acquire each preset pixel value, and perform pixel marking on each newly added annotation region of each step image corresponding to the same step image group based on the same preset pixel value, to obtain updated step images; establish an annotation association relationship between the step image group and the step subtext corresponding to the same task step; perform pixel marking on each step subtext based on a preset pixel value corresponding to the step image group having an annotation association relationship with the step subtext, and perform text updating on each task text based on the obtained updated step subtext.

[0014] According to another aspect of the present application, an employee point data processing system is provided, comprising: The acquisition module is configured to control the video acquisition unit and the voice acquisition unit of the mobile acquisition device to respectively perform video acquisition and voice acquisition on the speech operation of the offline speaker, to obtain a recorded video and a recorded voice. The extraction module is configured to perform semantic extraction on the recorded voice based on the employee names corresponding to each employee terminal, to obtain each task text including different step subtexts. The recognition module is configured to perform video splitting on the recorded video based on each task text, and perform segment recognition on each task segment obtained, to obtain each step image group including different step images. The acquisition module is configured to send the task text and each step image group corresponding to the same employee terminal to the employee terminal, and acquire step completion data uploaded by each employee terminal based on a project deadline. The update module is configured to perform point updating on the employee points of each employee terminal based on the upload quantity of each step completion data.

[0015] According to the scheme of the present application, the server can flexibly track the speech operation of the offline speaker by controlling the video and voice acquisition units of the mobile acquisition device, and accurately acquire complete recorded video and recorded voice. On this basis, the server can accurately extract the task text corresponding to each employee based on semantic extraction on the recorded voice by each employee terminal, and can facilitate the employee to quickly and clearly understand the specific task corresponding to the employee and improve the viewing efficiency. Then, the server will direct push the task text and the step image group to the corresponding employee terminal, and collect the step completion data according to the project deadline. The server will perform point updating on the employee points based on the upload quantity of the step completion data, which can quantify the work achievements of the employees, thereby effectively encouraging the employees to actively complete the corresponding tasks and further improving the overall work efficiency and project execution effect. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A flowchart of an employee point data processing method according to an embodiment of the present application is shown. Figure 2 A schematic diagram of an offset indicating line is shown according to an embodiment of the present application; Figure 3 A schematic diagram of a presentation area is shown according to an embodiment of the present application; Figure 4 A structural block diagram of an employee point data processing system according to another embodiment of the present application is shown. DETAILED DESCRIPTION

[0017] Exemplary embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is to be understood that the present disclosure can be embodied in various forms without being limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0018] To solve the problems in the background art, the inventors propose the present application. One embodiment of the present application provides an employee point data processing method, which can be executed in a computing device.

[0019] Figure 1 A flowchart of an employee point data processing method according to an embodiment of the present application is shown, which is suitable for execution in a computing device.

[0020] As Figure 1 shown, the employee point data processing method proposed in the present embodiment starts from step S102, in which the following contents are included: controlling a video acquisition unit and a voice acquisition unit of the mobile acquisition device to respectively acquire video and voice of the presentation operation of the offline presenter, to obtain recorded video and recorded voice.

[0021] For example, in the present embodiment, the mobile acquisition device can be flexibly adjusted in position to adapt to different offline presentation scenarios, such as large conference rooms, small discussion rooms, etc. When the offline presenter starts the presentation operation, the server controls the video acquisition unit and the voice acquisition unit of the mobile acquisition device to respectively acquire video and voice of the presentation operation of the offline presenter, so as to record the speech content and operation demonstration process of the offline presenter according to the obtained recorded video and recorded voice, to provide a complete data basis for subsequent task decomposition.

[0022] Further, the above-mentioned "controlling a video acquisition unit and a voice acquisition unit of the mobile acquisition device to respectively acquire video and voice of the presentation operation of the offline presenter, to obtain recorded video and recorded voice" further includes the following steps: When receiving a position debugging signal sent by the management end, the video acquisition unit of the mobile acquisition device is controlled to perform video acquisition to obtain a debugging video; Each video image frame constituting the debugging video is subjected to image comparison based on adjacent positions; When it is determined based on the comparison result that the two video image frames at the end have the same image content, the mobile acquisition device is controlled to end video acquisition, and any one of the two video image frames is determined as an identification reference frame; When it is determined through image identification of the identification reference frame that the position attribute corresponding to the identification reference frame is a suitable attribute, the video acquisition unit and the voice acquisition unit of the mobile acquisition device are controlled to perform video acquisition and voice acquisition on the speech operation of the offline speaker, respectively, to obtain a recording video and a recording voice.

[0023] For example, in the present embodiment, in order to ensure that the subsequent mobile acquisition device can completely and clearly acquire the operation demonstration process of the offline speaker, a suitable acquisition position can be debugged in the following manner: First, before starting the position debugging, the management end can send a position debugging signal. After receiving the position debugging signal, the server controls the video acquisition unit of the mobile acquisition device to perform video acquisition, thereby obtaining a debugging video. Then, the server performs image comparison based on adjacent positions on each video image frame constituting the debugging video. When the comparison result is that the two video image frames at the end have the same image content, it indicates that the device position of the mobile acquisition device has been stabilized. At this time, the server controls the mobile acquisition device to end video acquisition, and determines any one of the two video image frames as an identification reference frame. Finally, the server performs image identification on the identification reference frame to determine whether the position attribute corresponding to the identification reference frame is a suitable attribute. When it is determined that the position attribute corresponding to the identification reference frame is a suitable attribute, it indicates that the shooting angle and the picture integrity of the mobile acquisition device meet the acquisition requirements. Then, the server controls the video acquisition unit and the voice acquisition unit of the mobile acquisition device to perform video acquisition and voice acquisition on the speech operation of the offline speaker, respectively, thereby obtaining a recording video and a recording voice.

[0024] Furthermore, the above method further includes the following steps: The identification reference frame is subjected to binarization processing, and the obtained binarized image is subjected to pixel identification to obtain each image pixel point constituting the binarized image, wherein the binarized image includes a screen pixel point corresponding to a first pixel value and / or a noise pixel point corresponding to a second pixel value; If at least one image pixel point corresponding to the first pixel value is included in each image pixel point, the at least one image pixel point is determined as each edge pixel point, and whether each edge pixel point is located on the image contour corresponding to the binary image is determined. If none of the edge pixel points is located on the image contour of the binary image, it is determined that the position attribute corresponding to the identification reference frame is a suitable attribute.

[0025] For example, in the embodiment, the server performs binaryzation processing on the identification reference frame, thereby obtaining a binary image; Then, the server performs pixel identification on the binary image, thereby obtaining each image pixel point constituting the binary image. Since the mobile collection device is collecting the operation demonstration screen, the screen pixel point corresponding to the first pixel value and / or the noise pixel point corresponding to the second pixel value of the operation demonstration screen may be included in the binary image. If at least one image pixel point corresponding to the first pixel value is included in each image pixel point, the server determines the at least one image pixel point as each edge pixel point, and determines whether each edge pixel point is located on the image contour corresponding to the binary image. If none of the edge pixel points is located on the image contour of the binary image, it is determined that the position attribute corresponding to the identification reference frame is a suitable attribute.

[0026] Further, the above method further includes the following steps: If any edge pixel point is located on the image contour of the binary image, each image sub-contour constituting the image contour is determined. Each edge pixel point located on each image sub-contour and showing continuous adjacency is determined as a same frame contour segment, thereby obtaining each frame contour segment located in the identification reference frame corresponding to the binary image. The number of line segments corresponding to each frame contour segment is determined, and if the number of line segments is 1, the midpoint of the line segment corresponding to the frame contour segment is determined. Based on the identification reference frame, an offset indication line is generated, which has the midpoint of the line segment as a starting point and points to the image center point corresponding to the identification reference frame, and a debugging indication image is obtained and sent to the management end.

[0027] For example, in the embodiment, if any edge pixel point is located on the image contour of the binary image, it is determined that the video collection unit of the mobile collection device does not completely collect the operation demonstration screen. At this time, the server first determines each image sub-contour constituting the image contour. Then, the server determines each edge frame pixel point located in each image sub-contour and presenting continuous adjacency as a same frame contour segment, so as to obtain each frame contour segment located in the recognition reference frame corresponding to the binary image; Then, the server determines the number of line segments corresponding to each frame contour segment, if the number of line segments is 1, it indicates that the shooting angle of the mobile collection device exists a certain deviation in the direction opposite to the frame contour segment; Therefore, the server determines the midpoint of the line segment corresponding to the frame contour segment, and then generates a deviation indication line in the recognition reference frame, which has the midpoint as the starting point and points to the image center point of the corresponding recognition reference frame, as shown in Figure 2 Thus, the debugging indication image is obtained. Finally, the server sends the debugging indication image to the management end, so that the management end can perform position debugging on the mobile collection device according to the deviation indication line in the debugging indication image, and ensure the final collection effect.

[0028] In step S104, the following contents are included: Based on the employee names corresponding to each employee terminal, the semantic extraction of the recorded voice is performed to obtain each task text including different step subtexts.

[0029] For example, in the present embodiment, it can be understood that the offline speaker will arrange a work task including at least one task step for each employee according to the preset step template when performing the speaking operation, for example, "employee A completes data collection in the first step and performs preliminary analysis in the second step"; The server will perform semantic extraction on the recorded voice according to the employee names corresponding to each employee terminal, and split the recorded voice into task texts exclusive to each employee. In addition, the server will further split the task texts according to the step order corresponding to each task step to form step subtexts corresponding to each task step, which is more convenient for each employee to quickly view his own work task based on the task text including different step subtexts corresponding to him.

[0030] Further, the above "based on the employee names corresponding to each employee terminal, the semantic extraction of the recorded voice is performed to obtain each task text including different step subtexts" further includes the following steps: The preset participant list is called, wherein the preset participant list includes each employee terminal and the employee names corresponding to each employee terminal; The semantic recognition of the recorded voice is performed, and the obtained voice text is segmented based on the employee names of the preset participant list to obtain each task text corresponding to each employee terminal corresponding to each employee name; The server labels each task text based on the step keywords corresponding to different preset step templates, to obtain each step subtext corresponding to different task steps.

[0031] For example, in this embodiment, the server first calls the preset attendance list, which includes each employee terminal that should participate in the current presentation meeting and the employee name corresponding to each employee terminal. The employee name can be understood as the employee's name; Then, the server performs semantic recognition on the recorded voice to obtain a voice text, and performs text segmentation on the voice text according to the employee name in the preset attendance list, to obtain the task text corresponding to each employee terminal corresponding to each employee name; Finally, the server labels each task text based on the step keywords corresponding to different preset step templates, for example, the step keywords can be: first step, second step, etc., so as to divide the text based on the step label of each task text, to obtain each step subtext corresponding to different task steps.

[0032] In step S106, the following contents are included: The server performs video splitting on the recorded video based on each task text, and performs segment recognition on each task segment obtained, to obtain each step image group including different step images.

[0033] For example, in this embodiment, since each task text corresponds to a specific time period in the presentation operation, the server can perform video splitting on the recorded video based on each task text, to obtain a task segment corresponding to each employee terminal; Then, the server performs further segment recognition on the task segment, to obtain each step image group including different step images. The step image can be understood as a demonstration image in the operation demonstration screen of the offline presenter when corresponding to the presentation of a certain task step. That is, the employee terminal can quickly view the step image corresponding to each task step in the step image group, to facilitate the employee to understand the corresponding task requirements in combination with the step image.

[0034] Further, the above-mentioned "performing video splitting on the recorded video based on each task text, and performing segment recognition on each task segment obtained, to obtain each step image group including different step images" further includes the following steps: The server performs video splitting on the recorded video based on each task text, and performs segment recognition on each task segment obtained, to obtain each step image group including different step images, including: The server performs video splitting on the recorded video based on the obtained start time and end time corresponding to each task text, to obtain each task segment corresponding to each employee terminal; split the video according to the step time corresponding to each step mark to obtain each step segment corresponding to each step subtext; image comparison is performed on each video image frame corresponding to the same step segment in sequence, and a region change image is determined for each video image frame with a difference region based on the corresponding comparison result; Each region change image corresponding to the same step segment is sequentially filled with a serial number in a preset mark region based on the time sequence, and each step image corresponding to the same step segment is divided into the same step image group.

[0035] For example, in this embodiment, the server first acquires the start time and end time corresponding to each task text, and then performs video splitting on the recorded video according to the start time and end time, thereby obtaining each task segment corresponding to each employee terminal; Then, the server acquires the step time corresponding to each step mark, and then performs video splitting on each task segment according to the step time, thereby obtaining each step segment corresponding to each step subtext; Then, the server performs image comparison on each video image frame corresponding to the same step segment in sequence to determine whether there is a difference region between each video image frame; Since the operation demonstration screen can be large, the offline presenter can demonstrate from one area of the operation demonstration screen to another area during presentation. In order to avoid the situation that the demonstration traces in the step image generated subsequently are too many or overlapping, so that the employee cannot clearly determine which demonstration trace is associated with the step image; The server determines a region change image for each video image frame with a difference region based on the corresponding comparison result, and sequentially fills each region change image corresponding to the same step segment with a serial number in a preset mark region based on the time sequence, thereby obtaining each step image; Finally, the server divides each step image corresponding to the same step segment into the same step image group, so that the employee can clearly view the demonstration operation according to the image sequence based on the step image group subsequently.

[0036] Furthermore, the above-mentioned "image comparison is performed on each video image frame corresponding to the same step segment in sequence, and a region change image is determined for each video image frame with a difference region based on the corresponding comparison result" further includes the following steps: Each video image frame is divided into a left upper sub-region, a left lower sub-region, a right upper sub-region, and a right lower sub-region based on the image center point of each video image frame; The video image frames corresponding to the same step segment are divided into the same image determination group, and the video image frames in each image determination group are sequentially compared with the video image frames in front of the video image frames to obtain each difference region; Each sub-region corresponding to the difference region is determined as the presentation region corresponding to the video image frame; The video image frames corresponding to the same presentation region are divided into the same region image group according to the adjacent relationship, and the video image frames at the end of each region image group are determined as the region change image.

[0037] For example, in the embodiment, since the operation demonstration screen can be large, the server divides the video image frames into the upper left sub-region, the lower left sub-region, the upper right sub-region and the lower right sub-region according to the image center point of the video image frames; Then, the server divides the video image frames corresponding to the same step segment into the same image determination group, and sequentially compares the video image frames in each image determination group with the video image frames in front of the video image frames; If the difference region with image difference is obtained based on the image comparison, that is, the difference region is the region where the offline presenter performs the demonstration operation, therefore, the server determines each sub-region corresponding to the difference region as the presentation region corresponding to the video image frame, as shown in the difference region in the video image frame. Figure 3 The difference region in the upper left sub-region, therefore, the server determines the upper left sub-region as the presentation region corresponding to the video image frame; Finally, the server divides the video image frames corresponding to the same presentation region according to the adjacent relationship into the same region image group, that is, the presentation regions of all the video image frames in the region image group are the same, therefore, the server determines the video image frames at the end of each region image group as the region change image, that is, the region change image can completely present the final demonstration state of the presentation region, so as to facilitate the staff to quickly grasp the key task information.

[0038] Further, the above-mentioned "sequentially comparing the video image frames in each image determination group with the video image frames in front of the video image frames to obtain each difference region" further includes the following steps: The video image frames in each image determination group are divided into the same image comparison group with the video image frames in front of the video image frames; binarize each video image frame in each image comparison group to obtain each first image and each second image corresponding to the same image comparison group, wherein any first image comprises first image pixel points corresponding to first pixel values, and any second image comprises second image pixel points corresponding to second pixel values; create a transparent fusion layer and sequentially stack the first images and the second images corresponding to the same image comparison group in the transparent fusion layer; if any second image pixel point is included in the transparent fusion layer, perform pixel connection on the second image pixel points to obtain each difference region.

[0039] For example, in the embodiment, the server divides each video image frame in each step image group and the video image frame located in front of each video image frame into the same image comparison group; binarize each video image frame in each image comparison group to obtain each first image and each second image corresponding to the same image comparison group, wherein each first image comprises first image pixel points corresponding to first pixel values, and each second image comprises second image pixel points corresponding to second pixel values; Then, the server creates a transparent fusion layer and sequentially stacks the first images and the second images corresponding to the same image comparison group in the transparent fusion layer. If any second image pixel point is included in the transparent fusion layer, it indicates that the offline presenter performs a new demonstration operation in the video image frame corresponding to the second image. At this time, the server performs pixel connection on the second image pixel points to obtain each difference region.

[0040] Further, the above method further comprises the following steps: determine the step image located at the end of each step image group as a comparison image frame, and sort each step image group based on time from early to late to obtain a comparison sequence; perform image comparison on each step image in each step image group and the comparison image frame corresponding to the step image group located in front of each step image group based on the comparison sequence, and determine each region part with image difference as each newly added annotation region corresponding to each step image according to the comparison result; obtain each preset pixel value, and perform pixel marking on each newly added annotation region of each step image corresponding to the same step image group based on the same preset pixel value to obtain updated step images; establish an annotation association relationship between the step image group corresponding to the same task step and the step subtext; The step subtext is pixel marked based on the preset pixel value corresponding to the step image group having the annotation association relationship with the step subtext, and the task text is text updated based on the updated step subtext.

[0041] For example, in the embodiment, the server determines the step image located at the end of each step image group as the comparison image frame, and sorts each step image group based on time from early to late to obtain a comparison sequence. Then, the server compares each step image located in each step image group with the comparison image frame corresponding to the step image group located in front of the step image group based on the comparison sequence, and determines each region part with image difference as each new annotation region corresponding to each step image. Then, the server obtains each preset pixel value, and pixel marks each new annotation region of each step image corresponding to the same step image group based on the same preset pixel value to obtain updated step images, so that the subsequent employee terminal can locate each demonstration mark corresponding to one task step according to the same preset pixel value when viewing. Then, the server establishes an annotation association relationship between the step image group corresponding to the same task step and the step subtext, and then pixel marks each step subtext based on the preset pixel value corresponding to the step image group having the annotation association relationship with the step subtext, to obtain updated step subtexts, and then text updates each task text according to the updated step subtext, so that the subsequent employee terminal can view each step image and step subtext corresponding to the same task step according to the same preset pixel value, to improve the viewing efficiency.

[0042] In step S108, the following is included: The task text and each step image group corresponding to the same employee terminal are sent to the employee terminal, and the step completion data uploaded by each employee terminal is obtained based on the project deadline.

[0043] For example, in the embodiment, the server sends the task text and each step image group corresponding to the same employee terminal to the employee terminal, so that each employee terminal can clearly view each task step and efficiently complete each task step. When the employee completes the corresponding task step based on the task text and the step image group, the step completion data corresponding to the task step needs to be uploaded, such as the checked table and the written document, and the server continuously receives and stores these step completion data before the project deadline.

[0044] In step S110, the following is included: The employee points of each employee terminal are updated based on the uploading quantity of each step completion data.

[0045] For example, in the present embodiment, the corresponding point rules can be preset in advance in the server, for example, 1 point is obtained for completing one task step; The server will count the uploading quantity of the step completion data uploaded by each employee terminal, and then update the employee points of each employee terminal according to the point rules.

[0046] According to the scheme of the present application, the server can flexibly track the speech operation of the offline speaker by controlling the video and voice collection units of the mobile collection device, and accurately obtain complete recording video and recording voice; on this basis, the server can extract the semantic of the recording voice based on each employee terminal, accurately extract the task text corresponding to each employee, split the recording video in combination with the task text, and identify the step image group, so as to facilitate the employee to quickly and clearly understand the specific task corresponding to him, and improve the viewing efficiency; then, the server will push the task text and the step image group to the corresponding employee terminal, and collect the step completion data according to the project deadline, and the server will update the employee points based on the uploading quantity of the step completion data, so as to quantify the employee's work results, thereby effectively encouraging the employee to actively complete the corresponding task, and further improving the overall work efficiency and project execution effect.

[0047] Another embodiment of the present application provides an employee point data processing system, Figure 4 The system includes: The collection module is configured to control the video collection unit and the voice collection unit of the mobile collection device to respectively collect the speech operation of the offline speaker to obtain recording video and recording voice; The extraction module is configured to extract the semantic of the recording voice based on the employee name of each employee terminal to obtain each task text including different step subtexts; The identification module is configured to split the recording video based on each task text, and identify each task fragment to obtain each step image group including different step images; The acquisition module is configured to send the task text and each step image group corresponding to the same employee terminal to the employee terminal, and acquire the step completion data uploaded by each employee terminal based on the project deadline; The update module is configured to update the employee points of each employee terminal based on the uploading quantity of each step completion data.

[0048] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order to not obscure the understanding of this description.

[0049] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order to not obscure the understanding of this description.

[0050] Similarly, it is to be understood that the mechanical details of the inventive features sometimes are grouped into a single embodiment, figure or description of related embodiments in this disclosure. However, this is done for the sake of brevity only, and it is not to be interpreted that the inventive features necessarily belong to the same embodiment.

[0051] It will be understood by those within the art that the modules, or units, or components of the devices in the examples disclosed herein can be arranged in a device differently than as described in the examples, or can be located in one or more devices that are different from those of the examples. The modules in the foregoing examples can be combined into a single module or further separated into multiple sub-modules.

[0052] It will be understood by those within the art that the modules, or units, or components of the devices in the examples can be adapted to be located in one or more devices that are different from those of the examples. The modules, or units, or components of the examples can be combined into a module, or unit, or component, and further can be divided into multiple sub-modules, or sub-units, or sub-components.

[0053] Furthermore, those skilled in the art will recognize that references to particular features contained in some embodiments of the application are not intended to imply that every embodiment of the application will include that particular feature. That is, the various aspects of the application can include one or more of the features described herein, and the various aspects of the application can be used independently of one another.

[0054] Furthermore, some of the embodiments described herein are of a "method" or a "process" that can be embodied in software, firmware or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of software, firmware or hardware. For a software implementation, the "process" or "method" can be implemented by one or more physical apparatus based on instructions that can be embodied on a computer readable medium such as the fixed (hard-coded) logic or the firmware stored in the memory of the data processor 100. Moreover, the software implemented application can be initialized by one or more of start up operations or threads.

[0055] As used herein, unless otherwise indicated, the use of the ordinal adjectives "first", "second", "third", etc., merely to distinguish different instances of a similar object do not imply a meaning that the objects must be in a given order, or that one comes before or after another.

[0056] While the application has been described in terms of several embodiments, those skilled in the art will recognize that the application can be practiced with modifications within the spirit and scope of the application, which are encompassed by the description. Furthermore, it is to be appreciated that the description set forth herein focuses on the functioning and utility of the application, and thus describes in some instances internal activities of a computer, data store, and hardware components in terms of operations performed by computer components. Such workflow descriptions are used by those skilled in the art to describe what pieces of hardware or software are involved in performing the described operations, and thus a particular piece of hardware or software should not be construed as being limited to performing only the operations described for that hardware or software.

Claims

1. A method for processing employee point-based data, characterized in that, include: The video acquisition unit and voice acquisition unit of the mobile acquisition device respectively capture video and voice of the offline speaker's speech operations to obtain recorded video and recorded voice. Semantic extraction is performed on the recorded audio based on the employee names corresponding to each employee's terminal, resulting in task texts that include sub-texts of different steps. Based on the text of each task, the recorded video is segmented, and the resulting task segments are segmented to obtain a group of images for each step, including images of different steps. Send the task text and image sets of each step corresponding to the same employee's terminal to the employee's terminal, and obtain the step completion data uploaded by each employee's terminal based on the project deadline; The employee points on each employee's end are updated based on the number of data uploaded in each of the aforementioned steps.

2. The method according to claim 1, characterized in that, The video acquisition unit and audio acquisition unit of the mobile acquisition device respectively capture video and audio of the offline speaker's presentation, resulting in recorded video and recorded audio, including: When the location debugging signal sent by the management terminal is received, the video acquisition unit of the mobile acquisition device is controlled to acquire video and obtain debugging video. Perform image comparison based on adjacent positions on each video image frame that makes up the debug video; When it is determined from the comparison results that the two video image frames at the end have the same image content, the mobile acquisition device is controlled to end the video acquisition, and either of the two video image frames is determined as the identification reference frame. When image recognition is performed on the recognition reference frame to determine that the position attribute corresponding to the recognition reference frame is a suitable attribute, the video acquisition unit and voice acquisition unit of the mobile acquisition device are controlled to perform video acquisition and voice acquisition of the offline speaker's speech operation, respectively, to obtain recorded video and recorded voice.

3. The method according to claim 2, characterized in that, The method further includes: The identification reference frame is binarized, and the resulting binarized image is pixel-recognized to obtain each image pixel that makes up the binarized image. The binarized image includes screen pixels corresponding to the first pixel value and / or noise pixels corresponding to the second pixel value. If each image pixel includes at least one image pixel corresponding to the first pixel value, the at least one image pixel is determined as each edge pixel, and it is determined whether each edge pixel is located on the image contour of the corresponding binarized image; If none of the edge pixels are located within the image contour of the binarized image, the position attribute corresponding to the recognition reference frame is determined to be an appropriate attribute.

4. The method according to claim 3, characterized in that, The method further includes: If any edge pixel exists within the image contour of the binarized image, determine the image sub-contours that make up the image contour; Each edge pixel that is located in each image sub-contour and is continuously adjacent is determined as the same bounding box contour segment, thus obtaining each bounding box contour segment located in the recognition reference frame corresponding to the binarized image. Determine the number of line segments corresponding to each border outline segment. If the number of line segments is 1, determine the midpoint of the line segment corresponding to the border outline segment. Based on the recognition reference frame, an offset indicator line is generated, starting from the midpoint of the line segment and pointing to the center point of the corresponding recognition reference frame, and the obtained debugging indicator image is sent to the management terminal.

5. The method according to claim 1, characterized in that, Semantic extraction is performed on the recorded audio based on the employee names corresponding to each employee's terminal, resulting in task texts that include sub-texts of different steps, including: Retrieve a preset list of participants, which includes each employee's terminal and the name of each employee corresponding to each employee's terminal; Semantic recognition is performed on the recorded audio, and the audio text is segmented based on the names of each employee in the preset meeting list to obtain the task text corresponding to each employee's name. Based on the keywords of each step corresponding to different preset step templates, the task text is marked with steps to obtain the sub-text of each step corresponding to different task steps.

6. The method according to claim 5, characterized in that, Based on the task text, the recorded video is segmented, and segment recognition is performed on the resulting task segments to obtain step image groups including images of different steps, including: Based on the start and end times of the corresponding task texts, the recorded video is split into segments to obtain task segments corresponding to each employee's terminal. Based on the obtained step times corresponding to each step marker, each task segment is video-splitting to obtain each step segment corresponding to each step subtext. The video image frames that are adjacent to each other in the same step segment are compared sequentially, and the regions that are different are determined and the images are changed based on the comparison results. The images of each region corresponding to the same step segment are filled with serial numbers in a preset marked area based on the chronological order, and the resulting images of each step corresponding to the same step segment are divided into the same step image group.

7. The method according to claim 6, characterized in that, The process involves sequentially comparing adjacent video image frames corresponding to the same step segment, and determining regions to modify based on the comparison results for each video image frame with differing regions. This includes: Based on the image center point of each video image frame, each video image frame is divided into the upper left sub-region, the lower left sub-region, the upper right sub-region, and the lower right sub-region. Each video image frame corresponding to the same step segment is divided into the same image determination group, and each video image frame in each image determination group is compared with the video image frame preceding each video image frame in turn to obtain each difference region; Each sub-region corresponding to each difference region is determined as the speech region corresponding to the video image frame; Video image frames that correspond to the same presentation area and have an adjacent relationship are divided into the same area image group, and the video image frames at the end of each area image group are determined as the area change images.

8. The method according to claim 7, characterized in that, Each video image frame in each image determination group is sequentially compared with the video image frame preceding it to obtain the difference regions, including: Each video image frame in each image determination group is assigned to the same image comparison group as the video image frame preceding each video image frame. Each video image frame located in each image comparison group is binarized to obtain each first image and each second image corresponding to the same image comparison group. Each first image includes each first image pixel corresponding to a first pixel value, and each second image includes each second image pixel corresponding to a second pixel value. Create a transparent blending layer, and then stack the first and second images corresponding to the same image comparison group onto the transparent blending layer in sequence; If the transparent blending layer includes any second image pixel, connect the pixels of the second image to obtain the different regions.

9. The method according to claim 7, characterized in that, The method further includes: The last step image in each step image group is determined as the comparison image frame, and the step image groups are sorted from first to last based on time to obtain the comparison sequence; Based on the comparison sequence, each step image in each step image group is compared with the comparison image frame corresponding to the step image group preceding each step image group. The regions where the comparison results show image differences are determined as the newly added annotation regions of each step image. Obtain each preset pixel value, and mark each newly added annotation area of ​​each step image in the same step image group with the same preset pixel value to obtain the updated step images; Establish the annotation association between step image groups and step sub-text corresponding to the same task step; Each step subtext is pixel-marked based on a preset pixel value corresponding to the step image group with which it has a labeling relationship, and the task text is updated based on the obtained updated step subtext.

10. An employee points-based data processing system, characterized in that, include: The acquisition module is configured to control the video acquisition unit and the voice acquisition unit of the mobile acquisition device to acquire video and voice of the offline speaker's speech operations, respectively, to obtain recorded video and recorded voice. The extraction module is configured to perform semantic extraction on the recorded audio based on the employee name of each employee terminal, and obtain task texts including sub-texts of different steps; The recognition module is configured to split the recorded video based on the text of each task, and to perform segment recognition on each of the obtained task segments to obtain a group of images of each step, including images of different steps. The acquisition module is configured to send the task text and the image group of each step corresponding to the same employee terminal to the employee terminal, and to acquire the step completion data uploaded by each employee terminal based on the project deadline. The update module is configured to update employee points on each employee's end based on the number of data uploaded in each of the steps described.

Citation Information

Patent Citations

  • An integrated production device for zinc alloy forming and cooling

    CN115178729B