← 返回文献列表

IT Literature Intelligence 终审版本 (VERIFIED) 论文编号: 077 | 原始基线: V0_ZCODE_BASELINE | 语义审核: needs_revision | 图表审核: pass


Fast recurrent processing via ventrolateral prefrontal cortex is needed by the primate ventral stream for robust core visual object recognition

Kar · Neuron · 2021 · Zotero itemID=885

这篇文章在猕猴身上做了一件此前没人做过的事:用药物可逆地沉默腹外侧前额叶皮层(ventrolateral prefrontal cortex, vlPFC),同时用皮层内阵列记录下颞叶皮层(inferior temporal cortex, IT)的群体活动并测行为,从而在小于 200 ms 的自然时间尺度上因果检验"IT 的晚期响应依赖经 vlPFC 的快速循环加工"这一假说。它值得读,是因为结果不止回答了"循环信号从哪来",还给出一个漂亮的方向性证据:关掉 vlPFC 后,IT 的活动和猴的行为反而更像纯前馈深度网络模型。

研究背景

核心物体识别(core object recognition)——在眼动中心约 8° 范围内快速辨认物体——由腹侧视觉流完成,IT 群体活动中线性可读出的物体身份信息是行为充分表征的主要候选。但当时最好的模型是一族基本纯前馈的深度卷积神经网络(deep convolutional neural network, DCNN),它们既无法完全预测猕猴逐张图片的行为难度模式(Rajalingham et al., 2018),也无法解释 IT 响应的晚期动态。作者团队此前的工作(Kar et al., 2019)为本研究提供了关键的"弹药":对 1320 张测试图片逐张估计物体求解时间(object solution time, OST),即 IT 群体线性解码达到该图猴行为水平所需的时间;一部分图片在 90–120 ms 的早期时相就被"解决"(early-solved),另一部分要到 150–180 ms(late-solved),提示后者依赖前馈之外的循环(recurrent)计算。缺口在于:循环信号究竟由哪些脑环路计算并传回 IT?候选包括腹侧流内部双向通路、IT 附近的周嗅皮层、杏仁核、纹状体,以及下游的 vlPFC——vlPFC 与 IT 有强双向解剖连接、含物体类别选择神经元(Freedman et al.),且此前损伤/冷却研究提示前额叶能调制 IT,但没有任何工作在逐图分辨率上区分"整体增益调制"与"特异的循环计算"。

研究思路

论文的逻辑是"靶向破坏":既然 Kar et al. (2019) 已经把哪些图片依赖晚期 IT 表征标定出来,那么如果某个环路节点对循环计算至关重要,沉默它应该带来两个特异后果——(1) IT 群体编码中线性可解码的物体身份信息在晚期时相消失或变差,而早期时相基本不受影响;(2) 行为缺陷集中在 late-solved 图片上,early-solved 图片几乎无碍。作者据此预先设定三个假设(图 3C):H0,vlPFC 与 IT 物体编码无关(两族图片均无变化);H1,vlPFC 只起整体调制作用(如唤醒、增益),两族图片缺陷相同;H2,vlPFC 是关键循环节点,late-solved 图片缺陷显著更大。之所以先选 vlPFC 下手,是因为它在 IT 下游、与 IT 双向连接、含类别选择神经元、其活动在 IT 前馈响应消失后仍维持并在 IT 中再次出现(提示自上而下反馈),而且"下游沉默"在实验上直接可行——muscimol 注入 vlPFC,同步记录同侧 IT。这一设计把"相关性证据"升级为"必要性证据",同时保留了逐图的分辨率,使 H1 与 H2 可分离。

方法

被试为两只成年雄性恒河猴(N,9 岁;B,5 岁)。每只猴在 IT 皮层植入两个 96 通道 Utah 阵列(猴 B 在左半球、猴 N 在右半球),从后前轴覆盖 IT 不同部位;以 20 kHz 采样记录,分析基于多单元活动(multi-unit activity),最终从 384 个电极中选出视觉驱动强、图像排序信度高的 153 个 IT 位点。vlPFC 注射靶点先由单电极普查确定——寻找有强视觉驱动和粗类别选择性的位点(共记录 15 个 vlPFC 位点,猴 B 7 个、猴 N 8 个),再用结构 MRI 确认其与 Freedman 等报道的类别选择区解剖一致:主沟(principal sulcus)的外侧与腹侧(图 2A)。实验以"天"为单位分两类 session 交替进行(注射日/非注射日,注射日后至少恢复 1 天,每类至少 10 个 session):每天依次做被动注视任务、二选一物体辨别任务、再一个被动注视任务;注射日在首个被动注视(计入无药对照)之后,向 vlPFC 分 5 个深度(间隔 0.5 mm)注射 muscimol(GABAA 受体激动剂),约 30 分钟后开始采集。行为任务为:注视中央点 300 ms 后呈现 100 ms 测试图(来自 1320 张图片集,含 10 个物体类别:熊、象、面孔、苹果、车、狗、椅、飞机、鸟、斑马,多为 POV-Ray 渲染的合成图叠加随机背景),空屏 100 ms 后出现目标与干扰物的标准视图供选择,猴扫视到所选物体并保持注视 400 ms;每张图片的行为指标按 one-versus-all 方式对全部 9 个干扰物平均。刺激与任务沿用了 Kar et al. (2019),从而可直接使用其逐图 OST 标定:分析聚焦 OST 在 90–120 ms(early,209 张)与 150–180 ms(late,234 张)且行为 d' 在 2–4 之间的两组图片。神经分析的核心是逐 10 ms 时间窗训练线性支持向量机(support vector machine, SVM)解码器估计 IT 群体编码质量;模型比较方面,用偏最小二乘回归把前馈 DCNN 的倒数第二层("IT"层)特征映射到各 IT 位点,以逐位点预测力(predictivity,% 解释方差)的中位数评估匹配度,行为层面则比较逐图难度模式与模型的相关。

主要结果

  1. vlPFC 沉默选择性压低 IT 晚期平均响应(图 4B–D)。153 个 IT 位点上,早期时相(90–120 ms)平均响应无显著变化(配对 t 检验,t(152)=0.59,p=0.56,图注中给出的 DRearly ≈ −18%±46.4%),而晚期时相(150–180 ms)出现小幅但显著的下降(DRlate ≈ −31.83%±10.4%,t(152)=8.59,p<0.0001)。响应开始下降的时间(约 140 ms)恰与 vlPFC 位点的视觉响应潜伏期吻合(图 2D)。

  2. vlPFC 沉默特异性破坏 late-solved 图片的 IT 群体编码(图 4E–G)。以各图自己的 OST 为比较时点,late-solved 图片的线性解码准确率下降约 2%±0.61%,early-solved 仅 0.16%±0.53%(t(441)=2.41,p=0.0165);150–180 ms 区间两组曲线差异更大(t(441)=7.3,p<0.0001)。这一交互在按行为准确率分组(d' < 2、2–2.5、2.5–3、> 3)和 10 个辨别子任务上均成立(t(9)=1.97,p=0.04),且当物体位于记录阵列对侧视野时显著更强(猴 N:对侧 5.8% vs 同侧 2.01%;猴 B:对侧 2% vs 同侧 0.1%;置换检验 p<0.001)——同侧/对侧的不对称性也排除了"饱足或动机"等非特异解释。

  3. 行为缺陷同样集中在 late-solved 图片上(图 5)。muscimol 使二选一辨别总体准确率下降 6.03%±0.3%(t(859)=17.13,p<0.0001),其中 late-solved 图片下降 7.4%±0.5%、early-solved 仅 4.76%±0.45%(t(441)=2.40,p=0.0085);反应时总体延长 46.3±2.1 ms(t(858)=16.37,p<0.001),late-solved 图片延长更多(55±3.9 vs 34±4.19 ms,t(441)=2.05,p=0.04)。对侧视野的行为缺陷同样更强。

  4. vlPFC 沉默使腹侧流"退回"前馈模式(图 6)。IT 晚期响应(150–180 ms)对前馈模型 AlexNet fc7 的匹配度显著提高:中位 %EV 从 21.98% 升至 28.28%(t(152)=8.55,p<0.0001),早期响应无变化;晚期响应的图像排序与早期响应本身更相似。行为层面,vlPFC 沉默后猴的逐图难度模式与前馈模型行为预测的相关从 0.43 升至 0.56(置换检验 p<0.0001)。含局部循环的 CORnet-S 同样在沉默后更匹配(%ΔEV=2.41%±0.47%,行为预测力 +0.14),说明该模型虽比纯前馈网络更"脑似",但仍缺 vlPFC 这一循环节点。作者还验证了把 IT 活动连接到行为的线性"linking 模型"本身在给药前后不变,说明缺陷出在 IT 表征质量而非读出机制。

图注解读

图 1 · 方法动机:物体求解时间(OST)的定义

原文图注:Figure 1. Motivation (A) Estimation of object solution time (OST). For each image presentation (an example image of a bear is shown; 100 ms), we counted the number of multi-unit spike events (see STAR Methods for details) per site in nonoverlapping 10-ms windows after stimulus onset to construct a single population activity vector per time bin. These population vectors (image-evoked neural features) were then used to train and test cross-validated linear support vector machine decoders (D) separately per time bin. The decoder outputs per image (over time) were then used to perform a binary match to sample task (see STAR Methods) and obtain neural decode accuracies at each time bin. The time at which the neural decodes equal the primates' (pooled monkey) performance was then computed as the OST for that specific image. (B) Temporal evolution of linearly decodable object identity information in IT on an image-by-image basis.

这张图是理解全篇的钥匙。(A) 逐图流程: stimulus 呈现 100 ms 后,把每 10 ms 窗的 IT 多单位发放构造成群体活动向量,训练交叉验证的线性 SVM 解码器,解码准确率追上猴行为水平的时间点即该图的 OST。(B) 给出 1320 张图的 OST 直方图与示例曲线:蓝线两张是早解图片(解码快速达到行为水平),红线两张是晚解图片(需要更久)。灰线标注猴平均行为准确率(d' ≈ 2.5 水平)。它支撑引言与"研究思路"部分——"早解/晚解"不是事后分组,而是此前研究预先标定的、指向循环加工的逐图指标。

Figure 1

图 2 · vlPFC 注射靶点定位及其功能性质

原文图注:Figure 2. Approximate Location and Functional Properties of Injection Targets in vlPFC (A) The left panel shows the approximate anterior-posterior (AP) boundaries (black dashed lines) of the chamber that was placed over vlPFC. The green line denotes the location of the coronal section displayed on the right. The arrow refers to more anterior locations matching the other coronal MRI images. The right panels show structural MRI images of the approximate targets of the muscimol injections (SAR, superior arcuate sulcus; PS, principal sulcus; IAR, inferior arcuate sulcus). Injections were made lateral and ventral to the principal sulcus (indicated by the green patch). (B) Sample neural responses from a vlPFC site (averaged across 10 repetitions and 80 images) exhibiting visual drive. Error bars denote SEM across images.

(A) 用结构 MRI 显示注射舱与 muscimol 靶点:位于主沟外侧腹侧(绿色区块),冠以 SAR、PS、IAR 三条沟的解剖标记。(B) 示例 vlPFC 位点的平均响应时间 course,证明靶点有视觉驱动。(C)(原文图注此处截断)示例位点的粗类别选择性:按物体类别平均的响应曲线。(D) 15 个 vlPFC 位点(猴 B 7 个、猴 N 8 个)的响应潜伏期分布。这张图的作用是把"我们打药的地方就是有视觉响应、有类别选择的 vlPFC"落到实处;其潜伏期(图 2D)还与图 4B 中 IT 响应开始下降的时间相互印证,是"vlPFC 反馈在约 140 ms 后抵达 IT"这一时间链的关键一环。

Figure 2

图 3 · 实验设计与三个假设

原文图注:Figure 3. Experimental Setup and Hypotheses (A) Pharmacological inactivation of vlPFC (ipsilateral to the IT recording location) with simultaneous IT population recordings. (B) We divided the experiments into two different sessions, without (gray boxes) and with (green boxes) muscimol injections, conducted on consecutive days, with the exception of a passive viewing session before injections on the days with muscimol injections (see STAR Methods). We repeated each session in the same order after a minimum gap of 1 day (empty boxes). We completed at least 10 sessions for each condition type. (C) Hypothesized effects of vlPFC inactivation. One hypothesis (H0) is that the robustness of the IT object codes for core object recognition (~200 ms of processing) does not rely at all on vlPFC, which predicts no change in IT responses or behavior for both early- (blue bar) and late-solved (red bar) images. Another hypothesis (H1) is that vlPFC plays an overall modulatory role in ventral stream computations, which predicts deficits in IT population solution goodness and behavior that are equal for both groups of images. Finally, a third hypothesis (H2) is that vlPFC as a critical recurrence node in the brain circuitry for core objection, which predicts larger IT population solution deficits and larger behavioral deficits for late-solved images.

(A) 示意同侧组合:一边向 vlPFC 注射 muscimol,一边在 IT 用 Utah 阵列同步记录。(B) 时间线:注射/非注射日交替,注射日前先做一次被动注视(作为当日内的对照),每类条件至少 10 个 session。(C) 用蓝/红柱对比早解与晚解图片在三个假设下的预期缺陷模式:H0 无变化、H1 等量缺陷(整体调制)、H2 晚解图片缺陷更大(关键循环节点),并注明 H1+H2 混合的可能。这张图是全篇的推理框架:后面所有结果都对照这三个假设读——最终数据落在 H2 为主、混有少量 H1 成分的位置。

Figure 3

图 4 · 神经结果:晚期 IT 群体编码被特异性削弱

原文图注:Figure 4. Neural Experiments and Results (A) We measured neural responses from 153 sites in the IT cortex across two monkeys while they performed a battery of core recognition tasks, with and without muscimol injections in vlPFC. (B) Normalized mean IT firing rate in the two conditions (black, no-muscimol control condition; green, after muscimol injections in vlPFC). The shades indicate SEM across images. (C) We observed no significant differences across neurons at the early phase (90–120 ms) of the IT responses (DRearly = −18% ± 46.4%, mean ± SEM; paired t test; t(152) = 0.5885, p = 0.5571). (D) We observed a small but significant reduction in firing rates at the late phase (150–180 ms) of the IT responses (DRlate = −31.83% ± 10.4%, mean ± SEM; paired t test; t(152) = 8.5906, p < 0.0001). Error bars for (C) and (D) denote the standard deviation of responses across images per neuron. (E) Images (n = 234) with late OST (in red) showed a significantly higher drop in IT population decode accuracy across time (left panel) and at their corresponding OST (right panel) upon vlPFC inactivation compared to the images (n = 209) with early OST (in blue). This comparison was made with all images that had a measured (behavioral) d0 between 2 and 4, as measured in separate animals (Kar et al., 2019). Error bars denote SEM across images. We quantified the strength of this interaction as the difference in the muscimol-induced change (right panel), and we refer to that measure as dIT.

(B) 黑线(对照)与绿线(给药)的平均 IT 发放率时间 course:两条线在约 140 ms 前几乎重合,之后绿线低下去——这是"晚期依赖 vlPFC"最直观的形态。(C)(D) 分别对早期和晚期时相做逐神经元配对比较:早期不显著、晚期显著下降。(E) 左图是逐 10 ms 的解码准确率变化,红(晚解)蓝(早解)两条曲线在 150–180 ms 分开;右图聚焦各图自己的 OST,定义交互量 dIT。(F) 显示 dIT 在不同行为准确率分组下恒为负(即晚解图片解码缺陷更大)。(G) 按对侧/同侧视野分组,交互在记录阵列对侧视野的图片上显著更强。此图支撑"主要结果"第 1、2 条。

Figure 4

图 5 · 行为结果:晚解图片的行为缺陷更大

原文图注:Figure 5. Behavioral Experiments and Results (A) We tested behavioral performance on ten object categories, where performance was derived from the corresponding 45 binary object discrimination tasks with those 10 categories. (B) Two example trials of the binary object discrimination task showing the timeline of events. Monkeys fixate on a central dot, and then the test image at 8° containing 1 of 10 possible objects is shown for 100 ms (shown is a car [left trial] and a zebra [right trial]). After a 100-ms delay, a canonical view of the target object and a distractor object (one of the other nine objects) appears (randomly assigned on each trial to the left and right positions), and the monkey indicates which object was present in the test image by making a saccade to one of the two choices. We compared performance on sessions with and without muscimol injections in vlPFC. (C) vlPFC inactivation resulted in a larger performance drop among images (n = 234) with late OST (red bar), compared to the images (n = 209) with early OST (see Results for statistics).

(A) 10 个物体类别及其两两组成的 45 个二选一任务。(B) 试次时间线:100 ms 测试图 → 100 ms 空屏 → 目标+干扰物选择屏,扫视选择。(C) 核心柱状图:late-solved(红)与 early-solved 图片的给药后成绩下降对比,晚解组下降更大(7.4% vs 4.76%)。(D)(图注此处截断)显示 dB 在不同行为准确率分组下恒小于零,即在大多数子任务上晚解图片缺陷更重。(E)(图注此处截断)对侧/同侧视野的行为缺陷比较,与图 4G 的神经结果同构。此图支撑"主要结果"第 3 条,并与图 4 形成"神经编码缺陷—行为缺陷"的对应。

Figure 5

图 6 · 与计算模型比较:沉默 vlPFC 使腹侧流更像前馈模型

原文图注:Figure 6. Comparison with Computational Models: vlPFC Inactivation Causes the Ventral Stream to Behave More Like Feedforward Models (A) We showed 683 images to the monkey (fixated passive viewing) while recording simultaneously from their IT cortex, with and without vlPFC inactivation (top panel). The dashed red line denotes the recurrent pathway between vlPFC and the primate ventral stream. We compared the IT responses with and without vlPFC inactivation to those of the penultimate ("IT") layers of a feedforward DCNN model of the ventral stream (bottom panel) using previously established methods (see STAR Methods). We also compared the pattern of monkeys' behavioral responses (pattern of difficulty over images, see Rajalingham et al., 2018) with and without vlPFC inactivation to the model's behavioral pattern. In both types of comparisons, the key measure is referred to as predictivity, as it assesses the goodness of model predictions on new images.

(A) 示意实验逻辑:683 张图被动注视下记录 IT,虚线红箭头标出 vlPFC 与腹侧流之间的循环通路;把给药前后的 IT 响应和猴行为分别与前馈 DCNN 的"IT"层及模型行为比较,指标是留出图片上的预测力(predictivity)。(B) 上排:AlexNet fc7 对 IT 早期/晚期响应的 %EV,给药前(黑)后(绿)对比——晚期显著升高、早期不变;下排:猴逐图难度模式与模型行为的相关从 0.43 升至 0.56。这张图支撑"主要结果"第 4 条,也是全文最反直觉的发现:去掉一个脑区,神经表征反而"更像"人工模型——因为正常脑比纯前馈模型多做的正是 vlPFC 那部分循环加工。

Figure 6

讨论

作者的解读是:vlPFC 并不是对 IT 做"简单增益调制"(H1),而是特异性地把 IT 分布式群体编码的"格式"向行为充分的解推进,这种改进集中在晚期时相(H2);移除 vlPFC 后腹侧流退化成一个浅层前馈系统,其行为缺陷恰落在对浅层前馈计算机视觉系统最难的那批图片上。这一结果与既往因果研究一致——冷却 MT 使 V1/V2 响应下降 20%–40%(Hupé et al., 1998)、前额叶损伤影响 IT 色彩选择性与对侧视野信息传递(Fuster et al., 1985;Tomita et al., 1999)——但作者强调,本研究首次在小于 200 ms 的自然时间尺度上、以同步大规模神经与行为测量、按逐图分辨率因果检验了 vlPFC 的必要性。讨论中认真处理了一个反例:Minamimoto et al. (2010) 发现切除双侧外侧前额叶的猴仍能学习并泛化知觉类别,作者以三点调和——单侧小范围可逆失活对大面积永久切除;该研究用的规范视角、无背景、无尺寸位置变化的简单图片类似本研究的 early-solved 图片;其呈现时长 0.5–1.5 s 且有预提示,远超核心识别的 200 ms 窗口。作者承认的局限包括:无法指明信息从 vlPFC 重入腹侧流的确切环路(前馈投射主要止于 45A/B、46v、12r/l,反馈散布于 TEO/TE,不排除经额叶眼区等间接通路);linking 模型层面"vlPFC 直接驱动行为"与"vlPFC 把计算产物送回 IT、经尾状核等驱动行为"两种假说无法裁决,而早解图片也出现较小缺陷提示两种成分可能并存;实验设计与分析也未提供 vlPFC 注射对 IT 之外脑区的扩散控制细节。文末提出的路线图是:把这些环路假设做成含 vlPFC 节点的图像可计算循环网络,用本研究发布的给药前后 IT 响应与逐图行为数据约束并淘汰模型。

一句话总结

我认为这篇论文的真正亮点是那个"反直觉"的比较:沉默 vlPFC 让 IT 和猴行为向纯前馈 DCNN 靠拢,等于用因果操作证明了"脑比模型强的部分,就在这条反馈环路上"。当然单侧 0.4 cm³ 量级的局部失活只动了一部分 vlPFC,缺陷也就百分之几——这既是谨慎,也提醒我们循环加工的贡献是"精修"而非"地基"(这是我的理解,非论文原话)。


审校与证据追溯 (Verification & Evidence)

图表审计结果

关键事实与局限性声明