|
JCSE, vol. 20, no. 3, pp.142-153, September, 2026
DOI: http://dx.doi.org/10.5626/JCSE.2026.20.3.142
Algorithm Optimization for Real-Time 3D Face Driving and Expression Reconstruction in Immersive Virtual Live Streaming
Yonghui Wang School of Media, Nanchang Institute of Technology, Nanchang, China
Abstract: Real-time 3D face driving and expression reconstruction are the core technologies of immersive virtual live-streaming. Existing methods suffer from insufficient high-dynamic expression capture accuracy, weak temporal coherence, and high end-to-end inference latency. This paper proposes a joint optimization framework that uses the 3D morphable model (3DMM) as its geometric basis and integrates a temporal-aware graph convolutional network. The framework introduces a multi-branch decoupled estimation structure with an orthogonal regularization term that reduces strong-expression MAE by 27?30% relative to detailed expression capture and animation (DECA) and Deep3DFace baselines. A temporal consistency loss, jointly backpropagated through a GRU temporal unit and a lightweight spatial encoder, reduces interframe jitter NME by 53.7% relative to a non-temporal baseline. A semantic adaptive weight mechanism, grounded in Facial Action Coding System-based facial region partition, reduces periocular and lip reconstruction errors by 34.0% and 37.7%, respectively. Engineering deployment via INT8 quantization and TensorRT operator fusion compresses end-toend inference latency from 23.9 ms to 8.9 ms, achieving a total system latency of 17.7 ms?47% below the 30 FPS livestreaming threshold of 33.3 ms?with an identity preservation score (ArcFace cosine similarity) of 0.91 on a 500-clip evaluation set covering 20 anchors across diverse demographics.
Keyword:
No keyword
Full Paper: 0 Downloads, 1 View
|