无双的技术博客 记录 AI、Linux、网络架构、FreeSWITCH 与企业数字化实践

所有标签

#career skills 1 #human verification 1 #workplace productivity 1 #AI collaboration 1 #职业能力 1 #提问能力 1 #人工核验 1 #职场效率 1 #生成式AI 1 #AI协作 1 #VoIP troubleshooting 0 #SIP trunk 0 #softphone calling 0 #SIP extensions 0 #FreeSWITCH tutorial 0 #VoIP 排障 0 #SIP 线路 0 #XML Dialplan 0 #软电话互拨 0 #SIP 分机 0 #Debian 12 0 #FreeSWITCH 教程 0 #open source 1 #business validation 1 #call control 1 #开源项目 1 #AI编程 1 #ESL 2 #业务验证 1 #呼叫控制 1 #FCC 2 #freeswitch教程 1 #inference economics 1 #AI side hustle 1 #concurrency tuning 1 #GPU colocation 1 #AI token business 1 #AI 数据中心 1 #并发调优 1 #GPU 服务器托管 1 #普通人创业 1 #卖 Token 1 #prompting 1 #website deployment 0 #beginner web development 0 #AI-assisted development 0 #PostgreSQL 1 #Vue 3 1 #Flask 1 #网站部署 1 #零基础建站 1 #AI 开发 1 #health 1 #remote work 1 #personal brand 1 #technical leadership 1 #business thinking 1 #technical career 1 #健康管理 1 #远程工作 1 #个人品牌 1 #技术管理 1 #业务思维 1 #技术职业发展 1 #fluid dynamics 0 #Boltzmann equation 0 #mathematics 0 #Hilbert's sixth problem 0 #Yu Deng 0 #Fields Medal 0 #流体力学 1 #波尔兹曼方程 1 #数学科普 1 #希尔伯特第六问题 1 #邓煜 1 #菲尔兹奖 1 #网络攻击 1 #ugging Face 0 #supply chain security 0 #sandbox escape 0 #cybersecurity 0 #供应链安全 1 #沙箱逃逸 1 #网络安全 0 #OpenAI 1 #Hugging Face 1 #团队信任 1 #管理责任 1 #向上沟通 1 #目标管理 1 #团队管理 1 #管理者沟通 1 #鲁迅风格 1 #职场写作 1 #风格迁移 1 #风格迁移 0 #Zsh 1 #Bash 1 #Shell 1 #Unix 1 #RTP语音 1 #TCP MSS 1 #报文分片 1 #GRE over IPSec 1 #钉钉机器人告警 1 #CPU 使用率 1 #Linux 资源监控 1 #Bash 服务器监控 1 #模型监控 1 #机器学习流水线 1 #Airflow 1 #数据预测 1 #XGBoost 1 #LLM evaluation 0 #acceptance criteria 0 #output format 0 #prompt engineering 0 #大模型评测 0 #验收标准 0 #输出格式 0 #Few-shot 0 #Prompt 工程 0 #CUDA compatibility 0 #model serving 0 #local inference 0 #CUDA 兼容 0 #模型服务 0 #本地部署 0 #RTX 3090 0 #Ollama 0 #quantization 0 #distillation 0 #mixture of experts 0 #dense model 0 #model selection 0 #量化 0 #模型蒸馏 0 #MoE 0 #Dense 0 #模型选型 0 #data validation 0 #table extraction 0 #vision model 0 #数据校验 0 #JSON Schema 0 #OCR 0 #表格识别 0 #视觉模型 0 #LLM security 1 #human approval 1 #authorization 1 #parameter validation 1 #tool use 1 #大模型安全 1 #人工确认 1 #权限控制 1 #参数校验 1 #工具调用 1 #Function Calling 2 #LLM applications 1 #data access 1 #conversation memory 1 #knowledge base 1 #web search 1 #大模型应用 1 #数据权限 1 #对话记忆 1 #知识库 1 #联网搜索 1 #logging 1 #Qwen 1 #Model Studio API 1 #日志 1 #API Key 2 #Python 3.10 2 #UV 2 #千问 1 #百炼 API 1 #generative AI 2 #context window 1 #large language models 1 #生成式 AI 1 #Top P 2 #Temperature 2 #上下文窗口 1 #Token 2 #大模型原理 1 #enterprise AI governance 0 #code review 0 #least privilege 0 #coding agent 0 #AGENTS.md 0 #企业 AI 治理 0 #代码审核 0 #最小权限 0 #编码 Agent 0 #AI Agent 1 #Kimi K3 0 #postmortem report 1 #复盘报告 1 #project management 1 #action items 1 #root cause analysis 1 #lessons learned 0 #AI-assisted writing 1 #project postmortem 1 #项目管理 1 #改进措施 1 #原因分析 1 #经验教训 0 #AI写作 1 #项目复盘 1 #AI cost governance 1 #open-weight models 1 #private AI deployment 1 #enterprise data security 1 #AI vendor risk 1 #AI成本治理 1 #开放权重模型 1 #私有化部署 1 #企业数据安全 1 #AI供应商风险 1 #Claude Design 2 #Anthropic 2 #ad disclosure 1 #authentic reviews 1 #content compliance 1 #product recommendation posts 1 #Xiaohongshu copywriting 1 #AI writing 1 #广告标识 1 #真实体验 1 #内容合规 1 #种草文案 1 #小红书文案 2 #AI 写作 1 #human review 1 #content editing 1 #text formatting 1 #AI copywriting 1 #ChatGPT 2 #人工复核 1 #内容整理 1 #文字排版 1 #AI 文案 2 #workflow 0 #content verification 0 #faithful formatting 0 #prompt design 3 #AI text formatting 0 #工作流 0 #内容校对 0 #Markdown 0 #保真排版 0 #提示词 5 #AI 文字排版 0 #LLM inference 1 #tokenizer 1 #model weights 1 #LLM inference 1 #Transformer 2 #分词器 1 #模型权重 1 #大模型推理 2 #PDF内容提炼 1 #提示词写作 1 #演示文稿设计 1 #AI生成PPT 1 #data security 1 #enterprise architecture 1 #process redesign 1 #operating governance 1 #enterprise management 1 #operating system 1 #data governance 1 #digital china 1 #数据安全 1 #企业架构 1 #流程再造 1 #运营治理 1 #企业管理 1 #经营系统 1 #数字中国 1 #workforce disruption 1 #business model 1 #cloud infrastructure 1 #enterprise transformation 1 #AI bubble 1 #Oracle 1 #岗位替代 1 #商业模式 1 #云计算 1 #企业数字化转型 1 #AI泡沫 1 #甲骨文 1 #concurrency benchmarking 1 #performance monitoring 1 #并发压测 1 #性能监控 1 #Prometheus 2 #TPOT 2 #TTFT 2 #LLM inference acceleration 1 #draft model 1 #acceptance rate 1 #大模型推理加速 1 #接受率 1 #EAGLE 2 #Medusa 2 #Speculative Decoding 2 #推测解码 1 #AI implementation 1 #hybrid retrieval 1 #private knowledge base 1 #knowledge governance 1 #enterprise RAG 1 #AI落地 1 #文档审核 1 #pgvector 1 #知识治理 1 #企业RAG 1 #LuminaRAG 2 #cross-functional collaboration 1 #organizational change 2 #business process reengineering 1 #CTO perspective 1 #enterprise digitalization 1 #digital transformation 2 #企业数字化 1 #CTO视角 1 #业务流程再造 1 #组织变革 1 #ERP迁移 1 #数据治理 2 #经营透明度 1 #跨部门协同 1 #数字化转型 2 #LLM serving 1 #GPU memory optimization 1 #显存优化 1 #高并发 AI 服务 1 #Continuous Batching 2 #PagedAttention 2 #KV Cache 3 #sft 2 #model evaluation 2 #data engineering 2 #LLM fine-tuning 2 #max_length 1 #OOM 1 #vLLM 7 #LoRA 1 #realm 1 #容器化 1 #docker 1 #Iptables 1 #VPN 1 #GRE 1 #eBPF 1 #Gre over IPSec 2 #MTU 2 #Freeswitch多注册 1 #中继网关 1 #Linux 1 #gprof 1 #perf 1 #BGP 1 #origin 1 #AS_PATH 1 #MED 1 #localpref 1 #AI Coding 2 #Claude Code 1 #Deepseek 1 #VOIP 3 #Freeswitch 4 #开源呼叫系统 1 #freeswitch模块 2 #fs_mod 1 #voip 1 #freeswitch 4 #Qdrant 1 #Milvus 1 #pgvector 1 #RAG 3 #Halo 0

How Weights Become Words: LLM Inference Explained

Open models include weights, tokenizers, and configuration—not a chat app. Follow one prompt through vLLM/Ollama: IDs become vectors, logits, text, then EOS. What makes it speak and stop?

Administrator Administrator 发布于 2026-07-16

破译大模型推理:从权重到文本

下载开源大模型后,磁盘里躺着的不是可执行聊天程序,而是权重、分词器和配置等静态文件。我沿着一次真实推理拆开它们如何被 vLLM/Ollama 驱动:文本怎样变成 token、向量与概率,又怎样流式返回;最后究竟是 EOS、停止规则还是 EOF 让模型闭嘴?这篇文章还会解释 KV Cache 为何能让生成显著提速。

Administrator Administrator 发布于 2026-07-16

What Metrics Matter for vLLM Monitoring and Tuning

vLLM tuning should not stop at checking whether the API returns an answer. TTFT, TPOT, throughput, KV Cache usage, queue length, and tail latency form the real feedback loop.

Administrator Administrator 发布于 2026-07-07

vLLM 性能监控要看哪些指标

vLLM 调优不能只看服务是否返回答案,而要围绕 TTFT、TPOT、吞吐量、KV Cache 使用率和尾延迟建立监控闭环。本文用本地脚本串起启动、测量、压测和指标采集。

Administrator Administrator 发布于 2026-07-07

KV Cache Is the Real Memory Killer in High-Concurrency AI Services

In high-concurrency LLM serving, model weights are rarely the only memory problem. KV Cache grows with requests and tokens, so PagedAttention and Continuous Batching are the real keys to vLLM throughput.

Administrator Administrator 发布于 2026-07-06

KV Cache 才是高并发 AI 服务的显存杀手

高并发大模型服务真正容易爆掉的不是模型参数,而是随请求增长的 KV Cache。理解 PagedAttention 和 Continuous Batching,才能看懂 vLLM 为什么能把吞吐做上去。

Administrator Administrator 发布于 2026-07-06

vLLM推理OOM排查记:不是显存不够,是你没搞清楚max_length和batch_size的坑

70B模型配了4张80G卡,长文本一推就爆。查了一圈发现不是显存容量问题,是max_length设太大,kv cache按最坏情况预分配显存。上线前没做profiling,差点多花20万买卡。

Administrator Administrator 发布于 2026-07-04