Stage video data processing method and system based on image processing

Through image processing technology, multi-angle cameras and neural networks are used to optimize motion capture data, solving the problems of motion capture device battery life and data accuracy, and achieving a more accurate stage virtual display effect.

CN120017773AInactive Publication Date: 2025-05-16GUANGZHOU LANDE ELECTRONICS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510197356.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When obtaining stage data, the motion capture device is limited by the performance personnel's clothing props that cannot wear too many sensors, resulting in missed motion capture and insufficient battery life. The data obtained by traditional cameras cannot directly generate accurate virtual characters or scene motion data.

Method used

Using image processing method, video images are acquired through at least two cameras of different angles, the motion model is predicted using the initial neural network model, and combined with the data of the motion capture unit for verification and optimization, the target neural network model is generated, and the virtual characters or special effects are finally displayed through the display unit.

Benefits of technology

Improves the battery life of the motion capture sensor while obtaining more accurate motion models to ensure synchronization and accuracy of virtual characters or special effects with real stage performances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017773A_ABST
    Figure CN120017773A_ABST
Patent Text Reader

Abstract

The invention provides a stage video data processing method and system based on image processing, and belongs to the technical field of image processing.The stage video data processing method comprises the steps that first, a first video image and a second video image are obtained based on an image sensing unit in a first video period, and a predicted first motion model is obtained according to the first video image and the second video image; and then the second motion model is acquired through the synchronously acquired motion capture unit, and the neural network model is optimized by using the second motion model, so that a relatively accurate target neural network model can be obtained in the next period, and the model can be adopted to process a video picture in the next period. A more accurate motion model can be obtained while the endurance of the motion capture sensor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a stage video data processing method and system based on image processing. Background Art

[0002] Stage virtual display is a technology that combines virtual reality (VR), augmented reality (AR), mixed reality (MR) and computer graphics (CG) technology with stage performances. It generates virtual scenes, virtual characters or special effects through real-time or pre-processed video data, and integrates them with real stage performances to create an immersive audio-visual experience.

[0003] In the process of acquiring stage data, motion capture equipment is generally used to capture the actors' movements on stage and generate virtual characters or special effects data for implementation. However, due to factors such as the performers' costumes and props, they cannot wear too many sensors, and in some cases some movements may be missed. At the same time, the motion capture sensor has a low battery life, and long-term data acquisition may result in insufficient battery life. The data acquired by traditional cameras cannot directly generate accurate motion data that can be applied to virtual characters or scenes. Summary of the invention

[0004] The embodiments of the present application provide a stage video data processing method and system based on image processing to improve the above-mentioned problems.

[0005] In order to achieve the above objectives, this application adopts the following technical solutions: In the first aspect, the present application proposes a stage video data processing method based on image processing, which is applicable to a stage video processing system. The stage video processing system includes an image sensing unit, a motion capture unit, a display unit and a controller. The image sensing unit includes at least two cameras set at different angles. The method includes: The controller obtains a first video image and a second video image based on an image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model, wherein the initial neural network model is used to output a predicted motion model according to the input video images taken at different angles; The controller obtains a first motion model based on an output result of the initial neural network model; The controller obtains a second motion model based on the motion capture unit in the first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model; The controller acquires a third video image and a fourth video image in a second video cycle, where the second video cycle is a continuous cycle after the first video cycle, and inputs the third video image and the fourth video image into a target neural network model, and acquires a third motion model according to an output result of the target neural network; The controller controls the display unit to perform display based on the third motion model during the second video period.

[0006] In combination with the first aspect, optionally, the controller obtains a second motion model based on a motion capture unit in a first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the initial neural network model is a target neural network model after optimization. The method includes: The controller fits the first motion model with the second motion model, and obtains a fourth motion model according to the fitting result; The controller controls the display unit to display the fourth motion model during the first video period.

[0007] In combination with the first aspect, optionally, the controller obtains a second motion model based on a motion capture unit in a first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model, including: The controller acquires a second motion model at a first data acquisition frequency based on the motion capture unit in a first video period; The method also includes: Comparing the first motion model with the second motion model, and obtaining an error value according to the comparison result; The error value is compared with a preset value. If the error value is greater than the preset value, the controller acquires a fifth motion model at a second data acquisition frequency based on the motion capture unit in a second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the second data acquisition frequency is greater than the first data acquisition frequency.

[0008] In combination with the first aspect, optionally, the method further includes: The error value is compared with a preset value. If the error value is smaller than the preset value, the controller acquires a fifth motion model at a third data acquisition frequency based on the motion capture unit in the second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the third data acquisition frequency is smaller than the first data acquisition frequency.

[0009] In combination with the first aspect, optionally, the controller obtains a first video image and a second video image based on an image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model, and the initial neural network model is used to output a predicted motion model according to the input video images taken at different angles, including: The controller extracts the m-th frame image and the n-th frame image from the first video image and the second video image respectively, and the corresponding acquisition time of the m-th frame image and the n-th frame image is the same; The controller extracts features from the m-th frame image and the n-th frame image, wherein the extraction method adopts a convolutional neural network; The controller performs data fusion on the extracted features, inputs the results of the data fusion into a recursive neural network and performs iterative calculations, and obtains a predicted motion model according to the results of the iterative calculations.

[0010] In combination with the first aspect, optionally, the controller obtains a second motion model based on a motion capture unit in a first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model, including: The controller inputs the second motion model into the initial neural network model and determines a loss function, wherein the loss function is an error of the three-dimensional motion vector; The controller optimizes the initial neural network based on the output result of the loss function and obtains the target neural network model, wherein, during the optimization process, the priority of eliminating the second motion model is lower than the priority of eliminating the first motion model.

[0011] In combination with the first aspect, optionally, the controller controls the display unit to display based on the third motion model in the second video period, including: The display unit acquires the third motion model and the preset display model, and performs skeleton binding between the preset display model and the third motion model; The display unit displays the preset display model after bone binding.

[0012] In the second aspect, the present application proposes a stage video data processing system based on image processing, the stage video processing system includes an image sensing unit, a motion capture unit, a display unit and a controller, the image sensing unit includes at least two cameras set at different angles, and the system is configured as follows: The controller obtains a first video image and a second video image based on an image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model, wherein the initial neural network model is used to output a predicted motion model according to the input video images taken at different angles; The controller obtains a first motion model based on an output result of the initial neural network model; The controller obtains a second motion model based on the motion capture unit in the first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model; The controller acquires a third video image and a fourth video image in a second video cycle, where the second video cycle is a continuous cycle after the first video cycle, and inputs the third video image and the fourth video image into a target neural network model, and acquires a third motion model according to an output result of the target neural network; The controller controls the display unit to perform display based on the third motion model during the second video period.

[0013] In conjunction with the second aspect, optionally, the system is configured as: The controller obtains a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the initial neural network model is the target neural network model after the optimization. The method includes: The controller fits the first motion model with the second motion model, and obtains a fourth motion model according to the fitting result; The controller controls the display unit to display the fourth motion model during the first video period.

[0014] In conjunction with the second aspect, optionally, the system is configured as: The controller obtains a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model, including: The controller acquires a second motion model at a first data acquisition frequency based on the motion capture unit in a first video period; The system is also configured to: Comparing the first motion model with the second motion model, and obtaining an error value according to the comparison result; The error value is compared with a preset value. If the error value is greater than the preset value, the controller acquires a fifth motion model at a second data acquisition frequency based on the motion capture unit in a second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the second data acquisition frequency is greater than the first data acquisition frequency.

[0015] In conjunction with the second aspect, optionally, the system is configured as: The error value is compared with a preset value. If the error value is smaller than the preset value, the controller acquires a fifth motion model at a third data acquisition frequency based on the motion capture unit in the second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the third data acquisition frequency is smaller than the first data acquisition frequency.

[0016] In conjunction with the second aspect, optionally, the system is configured as: The controller obtains a first video image and a second video image based on an image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model. The initial neural network model is used to output a predicted motion model according to the input video images taken at different angles, including: The controller extracts the m-th frame image and the n-th frame image from the first video image and the second video image respectively, and the corresponding acquisition time of the m-th frame image and the n-th frame image is the same; The controller extracts features from the m-th frame image and the n-th frame image, wherein the extraction method adopts a convolutional neural network; The controller performs data fusion on the extracted features, inputs the results of the data fusion into a recursive neural network and performs iterative calculations, and obtains a predicted motion model according to the results of the iterative calculations.

[0017] In conjunction with the second aspect, optionally, the system is configured as: The controller obtains a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model, including: The controller inputs the second motion model into the initial neural network model and determines a loss function, wherein the loss function is an error of the three-dimensional motion vector; The controller optimizes the initial neural network based on the output result of the loss function and obtains the target neural network model, wherein, during the optimization process, the priority of eliminating the second motion model is lower than the priority of eliminating the first motion model.

[0018] In conjunction with the second aspect, optionally, the system is configured as: The controller controls the display unit to display based on the third motion model in the second video period, including: The display unit acquires the third motion model and the preset display model, and performs skeleton binding between the preset display model and the third motion model; The display unit displays the preset display model after bone binding.

[0019] A third aspect of an embodiment of the present invention provides an electronic device, the electronic device comprising: At least one processor; and, a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method proposed in the first aspect of the embodiment of the present invention.

[0020] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method provided in the first aspect of the embodiment of the present invention.

[0021] In summary, the above method and device have the following technical effects: The embodiment of the present application proposes a method and system for processing stage video data based on image processing. First, in a first video cycle, a first video image and a second video image are obtained based on an image sensing unit, and a predicted first motion model is obtained according to the first video image and the second video image. Then, the second motion model is obtained by a motion capture unit obtained synchronously, and the neural network model is optimized using the second motion model. In the next cycle, a more accurate target neural network model can be obtained, and the model can be used to process the video image in the next cycle, so that a more accurate motion model can be obtained while improving the endurance of the motion capture sensor. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A flowchart of a stage video data processing method based on image processing proposed in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] The embodiment of the present application proposes a stage video data processing method based on image processing, which is applicable to a stage video processing system. The stage video processing system includes an image sensing unit, a motion capture unit, a display unit and a controller. The image sensing unit includes at least two cameras set at different angles. Depending on the equipment and the venue, three or four cameras can also be set. The number of cameras is not limited in this application. The method includes steps S101-S105: S101: The controller obtains the first video image and the second video image based on the image sensing unit in the first video cycle, and inputs the first video image and the second video image into an initial neural network model, which is used to output a predicted motion model based on the input video images taken from different angles.

[0025] S102: The controller obtains a first motion model based on the output result of the initial neural network model.

[0026] It is understandable that in this embodiment, the specific time length of the first video cycle can be one frame or ten frames, which is not limited in this application. In the first video cycle, the camera obtains the first video image and the second video image respectively, and each image has multiple frames. The controller can output a predicted motion model based on the two video images at different angles. In this embodiment, the motion model is a three-dimensional motion vector model.

[0027] It can be understood that as an implementation method, a neural network model can be used to process data of a single video frame to obtain a predicted motion model, and then multiple motion models obtained according to different camera data are fitted to obtain the final motion model. Such a motion model has higher accuracy. Of course, it can also be obtained by using two images through feature recognition.

[0028] Specifically, in this embodiment, the controller may extract the mth frame image and the nth frame image from the first video image and the second video image respectively, and the corresponding acquisition time of the mth frame image and the nth frame image is the same, that is, they have the same timestamp.

[0029] Afterwards, the controller can perform feature extraction on the still m-th frame image and the n-th frame image, wherein the extraction method uses a convolutional neural network. A two-stream network (Siamese Network) with shared weights can be used to process the two images, or separate convolutional neural networks can be used to extract features respectively, which is not limited in this application.

[0030] After obtaining the extracted features, the controller performs data fusion on the extracted features. Common methods include concatenation, addition, or using an attention mechanism.

[0031] Afterwards, the result of data fusion can be input into a recursive neural network or a fully connected layer and iteratively calculated, and the predicted motion model can be obtained according to the result of the iterative calculation. The output motion model can be an optical flow field, a 3D motion vector, or other forms of motion representation. In this embodiment, a 3D motion vector is used as an example.

[0032] S103: The controller obtains the second motion model based on the motion capture unit in the first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the initial neural network model after optimization is the target neural network model.

[0033] It can be understood that in this embodiment, the motion capture unit may include sensors and receivers arranged on the stage characters. The sensors can capture the character's movements through data such as acceleration, combined with the wearing position. Therefore, the motion capture unit can obtain motion parameters that are more consistent with the actual situation. However, due to factors such as the costumes and props of the performers, it is impossible to wear too many sensors, and some actions may be missed in some cases. At the same time, the endurance of the motion capture sensor is low, and long-term data acquisition may result in insufficient endurance. Therefore, the data obtained by the motion capture unit can be combined with the data obtained by the neural network to obtain the final stage data.

[0034] It can be understood that in this embodiment, the controller inputs the second motion model into the initial neural network model and determines the loss function, wherein the loss function can be the error of the three-dimensional motion vector.

[0035] Then, the controller optimizes the initial neural network based on the output result of the loss function and obtains the target neural network model, wherein the priority of eliminating the second motion model during the optimization process is lower than the priority of eliminating the first motion model. It can be understood that since the motion capture unit basically conforms to the actual motion situation, the data obtained by the motion capture unit can be directly used as a verification set during the neural network training process to improve the accuracy of the output of the neural network model.

[0036] Optionally, the controller acquires the second motion model based on the motion capture unit at a first data acquisition frequency in the first video cycle. It can be understood that the higher the data acquisition frequency, the lower the endurance of the sensor, and vice versa.

[0037] S104: The controller acquires the third video image and the fourth video image in a second video cycle, where the second video cycle is a continuous cycle after the first video cycle, and inputs the third video image and the fourth video image into the target neural network model, and acquires the third motion model according to the output result of the target neural network.

[0038] It can be understood that in this embodiment, after the target neural network model is verified by the data obtained by the motion capture unit, its accuracy is higher than that of the initial neural network model, so the result output by the target neural network model can be used as the obtained result.

[0039] Optionally, the controller may compare the first motion model with the second motion model in the second cycle, and obtain an error value based on the comparison result. The larger the error value, the greater the difference between the model generated by the neural network and the actual model.

[0040] Then, the control can compare the error value with the preset value. If the error value is greater than the preset value, it proves that the deviation is large. Then the controller obtains the fifth motion model based on the motion capture unit at the second data acquisition frequency in the second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the second data acquisition frequency is greater than the first data acquisition frequency. In this way, the real model can be quickly obtained and the calculation process can be corrected.

[0041] Of course, if the error value is less than the preset value, the controller acquires the fifth motion model based on the motion capture unit at the third data acquisition frequency in the second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the third data acquisition frequency is less than the first data acquisition frequency. In this way, when the model generation is more accurate, the frequency of data acquisition can be reduced, and the endurance of the sensor can be improved.

[0042] S105: The controller controls the display unit to display based on the third motion model in the second video period.

[0043] It is understandable that in the present embodiment, the display of the display unit may be based on the third motion model. Exemplarily, the displayed content includes but is not limited to virtual characters or cartoon characters that are broadcast simultaneously. When displaying a virtual character or cartoon character, the display unit may first obtain the third motion model and a preset model of the virtual character or cartoon character, that is, a preset display model, and perform skeleton binding on the preset display model and the third motion model. The specific skeleton binding process has been disclosed in the relevant technical documents and is not limited in the present application. Then, the display unit displays the preset display model after skeleton binding, and the preset display model may make corresponding movements according to the motion model.

[0044] Optionally, for the model to be displayed in the first cycle, the controller may fit the first motion model with the second motion model, and obtain a fourth motion model according to the fitting result. Then, in the first video cycle, the preset display model is skeletally bound with the fourth motion model.

[0045] The present application proposes a stage video data processing method based on image processing. First, in a first video period, a first video image and a second video image are obtained based on an image sensing unit, and a predicted first motion model is obtained according to the first video image and the second video image. Then, a second motion model is obtained through a synchronously obtained motion capture unit, and a second motion model is obtained using The second motion model optimizes the neural network model, and a more accurate target neural network model can be obtained in the next cycle. The model can be used to process the video images in the next cycle, which can improve the endurance of the motion capture sensor and obtain a more accurate motion model.

[0046] Based on the same inventive concept, the embodiment of the present application also proposes a stage video data processing system based on image processing. The stage video processing system includes an image sensing unit, a motion capture unit, a display unit and a controller. The image sensing unit includes at least two cameras set at different angles. The system is configured as follows: The controller obtains a first video image and a second video image based on an image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model, wherein the initial neural network model is used to output a predicted motion model according to the input video images taken at different angles; The controller obtains a first motion model based on an output result of the initial neural network model; The controller obtains a second motion model based on the motion capture unit in the first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model; The controller acquires a third video image and a fourth video image in a second video cycle, where the second video cycle is a continuous cycle after the first video cycle, and inputs the third video image and the fourth video image into a target neural network model, and acquires a third motion model according to an output result of the target neural network; The controller controls the display unit to perform display based on the third motion model during the second video period.

[0047] Optionally, the system is configured to: The controller obtains a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the initial neural network model is the target neural network model after the optimization. The method includes: The controller fits the first motion model with the second motion model, and obtains a fourth motion model according to the fitting result; The controller controls the display unit to display the fourth motion model during the first video period.

[0048] Optionally, the system is configured to: The controller obtains a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model, including: The controller acquires a second motion model at a first data acquisition frequency based on the motion capture unit in a first video period; The system is also configured to: Comparing the first motion model with the second motion model, and obtaining an error value according to the comparison result; The error value is compared with a preset value. If the error value is greater than the preset value, the controller acquires a fifth motion model at a second data acquisition frequency based on the motion capture unit in a second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the second data acquisition frequency is greater than the first data acquisition frequency.

[0049] Optionally, the system is configured to: The error value is compared with a preset value. If the error value is smaller than the preset value, the controller acquires a fifth motion model at a third data acquisition frequency based on the motion capture unit in the second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the third data acquisition frequency is smaller than the first data acquisition frequency.

[0050] Optionally, the system is configured to: The controller obtains a first video image and a second video image based on an image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model. The initial neural network model is used to output a predicted motion model according to the input video images taken at different angles, including: The controller extracts the m-th frame image and the n-th frame image from the first video image and the second video image respectively, and the corresponding acquisition time of the m-th frame image and the n-th frame image is the same; The controller extracts features from the m-th frame image and the n-th frame image, wherein the extraction method adopts a convolutional neural network; The controller performs data fusion on the extracted features, inputs the results of the data fusion into a recursive neural network and performs iterative calculations, and obtains a predicted motion model according to the results of the iterative calculations.

[0051] Optionally, the system is configured to: The controller obtains a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model, including: The controller inputs the second motion model into the initial neural network model and determines a loss function, wherein the loss function is an error of the three-dimensional motion vector; The controller optimizes the initial neural network based on the output result of the loss function and obtains the target neural network model, wherein, during the optimization process, the priority of eliminating the second motion model is lower than the priority of eliminating the first motion model.

[0052] Optionally, the system is configured to: The controller controls the display unit to display based on the third motion model in the second video period, including: The display unit acquires the third motion model and the preset display model, and performs skeleton binding between the preset display model and the third motion model; The display unit displays the preset display model after bone binding.

[0053] The present application embodiment proposes a stage video data processing system based on image processing, firstly, in a first video cycle, a first video image and a second video image are obtained based on an image sensing unit, and a predicted first motion model is obtained according to the first video image and the second video image, and then a second motion model is obtained through a synchronously obtained motion capture unit, and a second motion model is obtained using The second motion model optimizes the neural network model, and a more accurate target neural network model can be obtained in the next cycle. The model can be used to process the video images in the next cycle, which can improve the endurance of the motion capture sensor and obtain a more accurate motion model.

[0054] Based on the same inventive concept, an embodiment of the present application further proposes an electronic device, the electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the stage video data processing method based on image processing according to an embodiment of the present application.

[0055] In addition, to achieve the above-mentioned purpose, an embodiment of the present application further proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the stage video data processing method based on image processing according to an embodiment of the present application.

[0056] The following is a detailed introduction to the various components of electronic equipment: The processor is the control center of the electronic device, which can be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs).

[0057] Optionally, the processor can perform various functions of the electronic device by running or executing a software program stored in the memory, and calling data stored in the memory.

[0058] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0059] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor, or may exist independently and be coupled to the processor through an interface circuit of the electronic device, which is not specifically limited in the embodiments of the present invention.

[0060] A transceiver is used to communicate with a network device or a terminal device.

[0061] Optionally, the transceiver may include a receiver and a transmitter, wherein the receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0062] Optionally, the transceiver may be integrated with the processor, or may exist independently and be coupled to the processor via an interface circuit of the router, which is not specifically limited in the embodiment of the present invention.

[0063] In addition, the technical effects of the electronic device can refer to the technical effects of the data transmission method in the above method embodiment, which will not be repeated here.

[0064] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0065] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0066] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0067] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0068] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0069] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0070] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

Claims

1. A stage video data processing method based on image processing, characterized in that: Applicable to a stage video processing system, the stage video processing system includes an image sensing unit, a motion capture unit, a display unit and a controller, the image sensing unit includes at least two cameras set at different angles, the method includes: The controller obtains a first video image and a second video image based on the image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model, wherein the initial neural network model is used to output a predicted motion model according to the input video images taken at different angles; The controller acquires a first motion model based on an output result of the initial neural network model; The controller acquires a second motion model based on the motion capture unit in the first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model; The controller acquires a third video image and a fourth video image in a second video cycle, where the second video cycle is a continuous cycle after the first video cycle, and inputs the third video image and the fourth video image into the target neural network model, and acquires a third motion model according to an output result of the target neural network; The controller controls the display unit to display based on the third motion model during the second video period.

2. The method for processing stage video data based on image processing according to claim 1, characterized in that: The controller acquires a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the initial neural network model is a target neural network model after optimization. The method includes: The controller fits the first motion model with the second motion model, and acquires a fourth motion model according to the fitting result; The controller controls the display unit to display the fourth motion model during the first video period.

3. The method for processing stage video data based on image processing according to claim 1, characterized in that: The controller acquires a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is a target neural network model, including: The controller acquires a second motion model at a first data acquisition frequency based on the motion capture unit during the first video period; The method further comprises: The controller compares the first motion model with the second motion model, and obtains an error value according to the comparison result; The controller compares the error value with a preset value. If the error value is greater than the preset value, the controller acquires a fifth motion model at a second data acquisition frequency based on the motion capture unit in the second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the second data acquisition frequency is greater than the first data acquisition frequency.

4. The method for processing stage video data based on image processing according to claim 3, characterized in that: The method further comprises: The error value is compared with a preset value. If the error value is smaller than the preset value, the controller acquires a fifth motion model at a third data acquisition frequency based on the motion capture unit in the second video cycle, and verifies and optimizes the result output by the neural network model based on the fifth motion model in the next continuous cycle, wherein the third data acquisition frequency is smaller than the first data acquisition frequency.

5. The method for processing stage video data based on image processing according to claim 1, characterized in that: The controller obtains a first video image and a second video image based on the image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model, wherein the initial neural network model is used to output a predicted motion model according to the input video images taken at different angles, including: The controller extracts an m-th frame image and an n-th frame image from the first video image and the second video image respectively, and the corresponding acquisition time of the m-th frame image and the n-th frame image is the same; The controller performs feature extraction on the m-th frame image and the n-th frame image, wherein the extraction method adopts a convolutional neural network; The controller performs data fusion on the extracted features, inputs the result of the data fusion into a recursive neural network and performs iterative calculation, and obtains the predicted motion model according to the result of the iterative calculation.

6. The method for processing stage video data based on image processing according to claim 5, characterized in that: The controller acquires a second motion model based on the motion capture unit in the first video cycle, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is a target neural network model, including: The controller inputs the second motion model into the initial neural network model and determines a loss function, wherein the loss function is an error of a three-dimensional motion vector; The controller optimizes the initial neural network based on the output result of the loss function and obtains the target neural network model, wherein during the optimization process, the priority of eliminating the second motion model is lower than the priority of eliminating the first motion model.

7. The method for processing stage video data based on image processing according to claim 1, characterized in that: The controller controls the display unit to display based on the third motion model in the second video period, including: The display unit acquires the third motion model and a preset display model, and performs skeleton binding between the preset display model and the third motion model; The display unit displays the preset display model after the skeleton binding.

8. A stage video data processing system based on image processing, characterized in that: The stage video processing system includes an image sensing unit, a motion capture unit, a display unit and a controller. The image sensing unit includes at least two cameras set at different angles. The system is configured as follows: The controller obtains a first video image and a second video image based on the image sensing unit in a first video period, and inputs the first video image and the second video image into an initial neural network model, wherein the initial neural network model is used to output a predicted motion model according to the input video images taken at different angles; The controller acquires a first motion model based on an output result of the initial neural network model; The controller acquires a second motion model based on the motion capture unit in the first video period, verifies the first motion model based on the second motion model and optimizes the algorithm of the initial motion model, and confirms that the model of the initial neural network model after optimization is the target neural network model; The controller acquires a third video image and a fourth video image in a second video cycle, where the second video cycle is a continuous cycle after the first video cycle, and inputs the third video image and the fourth video image into the target neural network model, and acquires a third motion model according to an output result of the target neural network; The controller controls the display unit to display based on the third motion model during the second video period.

9. An electronic device, characterized in that: Electronic equipment includes: at least one processor; and, a memory communicatively coupled to at least one of the processors; The memory stores instructions that can be executed by at least one of the processors, and the instructions are executed by at least one of the processors so that the at least one processor can execute the method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the method as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Speed precision optimization method and system of motion capture system based on neural network

    CN115861592A

  • Human motion capture method and device based on monocular RGB video and sparse IMU

    CN117911452A

  • Method and device with image processing function

    CN118057443A

  • Estimating camera pose

    US20220036577A1